paper-with-me

홈 › Papers

Predictive Speech Recognition and End-of-Utterance Detection Towards Spoken Dialog Systems

2024-09-30 · Oswald Zink, Yosuke Higuchi, Carlos Mullov, Alexander Waibel, Tetsunori Kobayashi

Effective spoken dialog systems should facilitate natural interactions with quick and rhythmic timing, mirroring human communication patterns. To reduce response times, previous efforts have focused on minimizing the latency in automatic speech recognition (ASR) to optimize system efficiency. However, this approach requires waiting for ASR to complete processing until a speaker has finished speaking, which limits the time available for natural language processing (NLP) to formulate accurate responses. As humans, we continuously anticipate and prepare responses even while the other party is still speaking. This allows us to respond appropriately without missing the optimal time to speak. In this work, as a pioneering study toward a conversational system that simulates such human anticipatory behavior, we aim to realize a function that can predict the forthcoming words and estimate the time remaining until the end of an utterance (EOU), using the middle portion of an utterance. To achieve this, we propose a training strategy for an encoder-decoder-based ASR system, which involves masking future segments of an utterance and prompting the decoder to predict the words in the masked audio. Additionally, we develop a cross-attention-based algorithm that incorporates both acoustic and linguistic information to accurately detect the EOU. The experimental results demonstrate the proposed model's ability to predict upcoming words and estimate future EOU events up to 300ms prior to the actual EOU. Moreover, the proposed training strategy exhibits general improvements in ASR performance.

📄 PDF Abstract BibTeX arXiv:2409.19990

Code (0)

등록된 구현이 없습니다.

Tasks

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Decoderspeech-recognitionSpeech Recognition

Similar Papers 제목 키워드 기반

Towards end-2-end learning for predicting behavior codes from spoken utterances in psychotherapy conversations

2020-07-01 · ACL 2020 6 · Karan Singla, Zhuohao Chen, David Atkins, Shrikanth Narayanan

Spoken language understanding tasks usually rely on pipelines involving complex processing blocks such as voice activity detection, speaker diarization and Automatic speech recognition (ASR). We propose a novel framework…

Action DetectionActivity DetectionAutomatic Speech RecognitionAutomatic Speech Recognition (ASR)+5

DeToxy: A Large-Scale Multimodal Dataset for Toxicity Classification in Spoken Utterances

2021-10-14 · Sreyan Ghosh, Samden Lepcha, S Sakshi, Rajiv Ratn Shah 외

Toxic speech, also known as hate speech, is regarded as one of the crucial issues plaguing online social media today. Most recent work on toxic speech detection is constrained to the modality of text and written conversa…

A Comparative Analysis of Crowdsourced Natural Language Corpora for Spoken Dialog Systems

2016-05-01 · LREC 2016 5 · Patricia Braunger, Hansj{\"o}rg Hofmann, Steffen Werner, Maria Schmidt

Recent spoken dialog systems have been able to recognize freely spoken user input in restricted domains thanks to statistical methods in the automatic speech recognition. These methods require a high number of natural la…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)speech-recognitionSpeech Recognition+1

Do We Still Need Automatic Speech Recognition for Spoken Language Understanding?

2021-11-29 · Lasse Borgholt, Jakob Drachmann Havtorn, Mostafa Abdou, Joakim Edin 외

Spoken language understanding (SLU) tasks are usually solved by first transcribing an utterance with automatic speech recognition (ASR) and then feeding the output to a text-based model. Recent advances in self-supervise…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Machine Translationnamed-entity-recognition+7

MaSS: A Large and Clean Multilingual Corpus of Sentence-aligned Spoken Utterances Extracted from the Bible

2019-07-30 · LREC 2020 5 · Marcely Zanon Boito, William N. Havard, Mahault Garnerin, Éric Le Ferrand 외

The CMU Wilderness Multilingual Speech Dataset (Black, 2019) is a newly published multilingual speech dataset based on recorded readings of the New Testament. It provides data to build Automatic Speech Recognition (ASR) …

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)RetrievalSentence+5