Evaluating OpenAI’s Whisper ASR for Punctuation Prediction and Topic Modeling of life histories of the Museum of the Person
-
Lucas Rafael Stefanel Gris
, Ricardo Marcacini , Arnaldo Candido Junior , Edresson Casanova , Anderson Soares and Sandra Maria Aluísio
Abstract
Automatic speech recognition (ASR) systems play a key role in applications involving human-machine interactions. Despite their importance, ASR models for the Portuguese language proposed in the last decade have limitations in relation to the correct identification of punctuation marks in automatic transcriptions, which hinder the use of transcriptions by other systems, models, and even by humans. However, recently, OpenAI proposed Whisper ASR, a general-purpose speech recognition model that has generated great expectations in dealing with such limitations. This chapter presents the first study on the performance of Whisper for punctuation prediction in the Portuguese language. We present an experimental evaluation considering both theoretical aspects involving pausing points (comma) and complete ideas (exclamation, question, and fullstop), as well as practical aspects involving transcript-based topic modeling – an application dependent on punctuation marks for promising performance. We analyzed experimental results from videos of Museum of the Person, a virtual museum that aims to tell and preserve people’s life histories, thus discussing the pros and cons of Whisper in a real-world scenario. Although our experiments indicate that Whisper achieves state-of-the-art results, we conclude that some punctuation marks require improvements, such as exclamation, semicolon and colon.
Abstract
Automatic speech recognition (ASR) systems play a key role in applications involving human-machine interactions. Despite their importance, ASR models for the Portuguese language proposed in the last decade have limitations in relation to the correct identification of punctuation marks in automatic transcriptions, which hinder the use of transcriptions by other systems, models, and even by humans. However, recently, OpenAI proposed Whisper ASR, a general-purpose speech recognition model that has generated great expectations in dealing with such limitations. This chapter presents the first study on the performance of Whisper for punctuation prediction in the Portuguese language. We present an experimental evaluation considering both theoretical aspects involving pausing points (comma) and complete ideas (exclamation, question, and fullstop), as well as practical aspects involving transcript-based topic modeling – an application dependent on punctuation marks for promising performance. We analyzed experimental results from videos of Museum of the Person, a virtual museum that aims to tell and preserve people’s life histories, thus discussing the pros and cons of Whisper in a real-world scenario. Although our experiments indicate that Whisper achieves state-of-the-art results, we conclude that some punctuation marks require improvements, such as exclamation, semicolon and colon.
Chapters in this book
- Frontmatter I
- Preface V
- Contents IX
- Prosody and L2 Learning Interface: The Case of Spanish L2 and Brazilian Portuguese L1 Intonation 1
- The Role of Prosody in the Processing of Ambiguities in Brazilian Portuguese 31
- Defining and Identifying Discourse Markers in Spontaneous Speech 65
- A Contribution to a Better Understanding of Silent Pause 103
- Perceptual and Physiological Correlates of Voice Quality Settings 127
- Multimodal Analysis of Speech Attractiveness Expression 151
- Posture and Gestures Can Affect the Prosodic Speaker Impact in a Remote Presentation 181
- An Acoustic Analysis of Creaky Voice Patterns in Singing 223
- Evaluating OpenAI’s Whisper ASR for Punctuation Prediction and Topic Modeling of life histories of the Museum of the Person 247
- Index
Chapters in this book
- Frontmatter I
- Preface V
- Contents IX
- Prosody and L2 Learning Interface: The Case of Spanish L2 and Brazilian Portuguese L1 Intonation 1
- The Role of Prosody in the Processing of Ambiguities in Brazilian Portuguese 31
- Defining and Identifying Discourse Markers in Spontaneous Speech 65
- A Contribution to a Better Understanding of Silent Pause 103
- Perceptual and Physiological Correlates of Voice Quality Settings 127
- Multimodal Analysis of Speech Attractiveness Expression 151
- Posture and Gestures Can Affect the Prosodic Speaker Impact in a Remote Presentation 181
- An Acoustic Analysis of Creaky Voice Patterns in Singing 223
- Evaluating OpenAI’s Whisper ASR for Punctuation Prediction and Topic Modeling of life histories of the Museum of the Person 247
- Index