paper-with-me

Papers

Conversational Speech Recognition Needs Data? Experiments with Austrian German

2022-06-01 · LREC 2022 6 · Julian Linke, Philip N. Garner, Gernot Kubin, Barbara Schuppler

Conversational speech represents one of the most complex of automatic speech recognition (ASR) tasks owing to the high inter-speaker variation in both pronunciation and conversational dynamics. Such complexity is particularly sensitive to low-resourced (LR) scenarios. Recent developments in self-supervision have allowed such scenarios to take advantage of large amounts of otherwise unrelated data. In this study, we characterise an (LR) Austrian German conversational task. We begin with a non-pre-trained baseline and show that fine-tuning of a model pre-trained using self-supervision leads to improvements consistent with those in the literature; this extends to cases where a lexicon and language model are included. We also show that the advantage of pre-training indeed arises from the larger database rather than the self-supervision. Further, by use of a leave-one-conversation out technique, we demonstrate that robustness problems remain with respect to inter-speaker and inter-conversation variation. This serves to guide where future research might best be focused in light of the current state-of-the-art.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Tasks

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Language ModelingLanguage Modellingspeech-recognitionSpeech Recognition

Similar Papers 제목 키워드 기반

Gated Embeddings in End-to-End Speech Recognition for Conversational-Context Fusion

2019-06-27 · ACL 2019 7 · Suyoun Kim, Siddharth Dalmia, Florian Metze

We present a novel conversational-context aware end-to-end speech recognizer based on a gated neural network that incorporates conversational-context/word/speech embeddings. Unlike conventional speech recognition models,…

SentenceSentence Embeddingsspeech-recognitionSpeech Recognition

Acoustic-to-Word Models with Conversational Context Information

2019-05-21 · NAACL 2019 6 · Suyoun Kim, Florian Metze

Conversational context information, higher-level knowledge that spans across sentences, can help to recognize a long conversation. However, existing speech recognition models are typically built at a sentence level, and …

Sentencespeech-recognitionSpeech Recognition

Conversational Speech Recognition By Learning Conversation-level Characteristics

2022-02-16 · Kun Wei, Yike Zhang, Sining Sun, Lei Xie 외

Conversational automatic speech recognition (ASR) is a task to recognize conversational speech including multiple speakers. Unlike sentence-level ASR, conversational ASR can naturally take advantages from specific charac…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)DecoderSentence+2

Improving Code-Switching Speech Recognition with TTS Data Augmentation

2026-01-02 · Yue Heng Yeo, Yuchen Hu, Shreyas Gopal, Yizhou Peng 외 arxiv

Automatic speech recognition (ASR) for conversational code-switching speech remains challenging due to the scarcity of realistic, high-quality labeled speech data. This paper explores multilingual text-to-speech (TTS) mo…

Speech RecognitionData Augmentation

The timing bottleneck: Why timing and overlap are mission-critical for conversational user interfaces, speech recognition and dialogue systems

2023-07-28 · Andreas Liesenfeld, Alianda Lopez, Mark Dingemanse

Speech recognition systems are a key intermediary in voice-driven human-computer interaction. Although speech recognition works well for pristine monologic audio, real-life use cases in open-ended interactive settings st…

Intent Recognitionspeech-recognitionSpeech Recognition