Towards Interactive Annotation for Hesitation in Conversational Speech
Manual annotation of speech corpora is expensive in both human resources and time. Furthermore, recognizing affects in spontaneous, non acted speech presents a challenge for humans and machines. The aim of the present study is to automatize the labeling of hesitant speech as a marker of expressed uncertainty. That is why, the NCCFr-corpus was manually annotated for {}degree of hesitation{'} on a continuous scale between -3 and 3 and the affective dimensions {}activation, valence and control{'}. In total, 5834 chunks of the NCCFr-corpus were manually annotated. Acoustic analyses were carried out based on these annotations. Furthermore, regression models were trained in order to allow automatic prediction of hesitation for speech chunks that do not have a manual annotation. Preliminary results show that the number of filled pauses as well as vowel duration increase with the degree of hesitation, and that automatic prediction of the hesitation degree reaches encouraging RMSE results of 1.6.
Code (0)
등록된 구현이 없습니다.
Similar Papers 제목 키워드 기반
Turn-Taking Prediction for Natural Conversational Speech
While a streaming voice assistant system has been used in many applications, this system typically focuses on unnatural, one-shot interactions assuming input from a single voice query without hesitation or disfluency. Ho…
Predictionspeech-recognitionSpeech RecognitionCurrent Challenges in Spoken Dialogue Systems and Why They Are Critical for Those Living with Dementia
Dialogue technologies such as Amazon's Alexa have the potential to transform the healthcare industry. However, current systems are not yet naturally interactive: they are often turn-based, have naive end-of-turn detectio…
Diagnosticspeech-recognitionSpeech RecognitionSpoken Dialogue SystemsHESITA(te) in Portuguese
Hesitations, so-called disfluencies, are a characteristic of spontaneous speech, playing a primary role in its structure, reflecting aspects of the language production and the management of inter-communication. In this p…
Acoustic ModellingAutomatic Speech RecognitionAutomatic Speech Recognition (ASR)Management+3Super-Human Performance in Online Low-latency Recognition of Conversational Speech
Achieving super-human performance in recognizing human speech has been a goal for several decades, as researchers have worked on increasingly challenging tasks. In the 1990's it was discovered, that conversational speech…
DecoderAcoustically Precise Hesitation Tagging Is Essential for End-to-End Verbatim Transcription Systems
Verbatim transcription for automatic speaking assessment demands accurate capture of disfluencies, crucial for downstream tasks like error analysis and feedback. However, many ASR systems discard or generalize hesitation…