Do We Still Need Automatic Speech Recognition for Spoken Language Understanding?
Spoken language understanding (SLU) tasks are usually solved by first transcribing an utterance with automatic speech recognition (ASR) and then feeding the output to a text-based model. Recent advances in self-supervised representation learning for speech data have focused on improving the ASR component. We investigate whether representation learning for speech has matured enough to replace ASR in SLU. We compare learned speech features from wav2vec 2.0, state-of-the-art ASR transcripts, and the ground truth text as input for a novel speech-based named entity recognition task, a cardiac arrest detection task on real-world emergency calls and two existing SLU benchmarks. We show that learned speech features are superior to ASR transcripts on three classification tasks. For machine translation, ASR transcripts are still the better choice. We highlight the intrinsic robustness of wav2vec 2.0 representations to out-of-vocabulary words as key to better performance.
Code (0)
등록된 구현이 없습니다.
Tasks
Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Machine Translationnamed-entity-recognitionNamed Entity RecognitionNamed Entity Recognition (NER)Representation Learningspeech-recognitionSpeech RecognitionSpoken Language UnderstandingTranslationSimilar Papers 제목 키워드 기반
SpeechDPR: End-to-End Spoken Passage Retrieval for Open-Domain Spoken Question Answering
Spoken Question Answering (SQA) is essential for machines to reply to user's question by finding the answer span within a given spoken passage. SQA has been previously achieved without ASR to avoid recognition errors and…
Passage RetrievalQuestion AnsweringRetrievalSentence+2Towards Data Distillation for End-to-end Spoken Conversational Question Answering
In spoken question answering, QA systems are designed to answer questions from contiguous text spans within the related speech transcripts. However, the most natural way that human seek or test their knowledge is via hum…
Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Conversational Question AnsweringQuestion Answering+2Automatic Spoken Language Identification Utilizing Acoustic and Phonetic Speech Information
Automatic spoken Language Identification (LID) is the process of identifying the language spoken within an utterance. The challenge that this task presents is that no prior information is available indicating the content…
Language Identificationspeech-recognitionSpeech RecognitionSpoken language identificationKnowledge Distillation for Improved Accuracy in Spoken Question Answering
Spoken question answering (SQA) is a challenging task that requires the machine to fully understand the complex spoken documents. Automatic speech recognition (ASR) plays a significant role in the development of QA syste…
Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Knowledge DistillationLanguage Modeling+5Improving Named Entity Recognition in Spoken Dialog Systems by Context and Speech Pattern Modeling
While named entity recognition (NER) from speech has been around as long as NER from written text has, the accuracy of NER from speech has generally been much lower than that of NER from text. The rise in popularity of s…
Automatic Speech RecognitionAutomatic Speech Recognition (ASR)named-entity-recognitionNamed Entity Recognition+4