paper-with-me

홈 › Papers

Label-Synchronous Neural Transducer for Adaptable Online E2E Speech Recognition

2023-11-19 · Keqi Deng, Philip C. Woodland

Although end-to-end (E2E) automatic speech recognition (ASR) has shown state-of-the-art recognition accuracy, it tends to be implicitly biased towards the training data distribution which can degrade generalisation. This paper proposes a label-synchronous neural transducer (LS-Transducer), which provides a natural approach to domain adaptation based on text-only data. The LS-Transducer extracts a label-level encoder representation before combining it with the prediction network output. Since blank tokens are no longer needed, the prediction network performs as a standard language model, which can be easily adapted using text-only data. An Auto-regressive Integrate-and-Fire (AIF) mechanism is proposed to generate the label-level encoder representation while retaining low latency operation that can be used for streaming. In addition, a streaming joint decoding method is designed to improve ASR accuracy while retaining synchronisation with AIF. Experiments show that compared to standard neural transducers, the proposed LS-Transducer gave a 12.9% relative WER reduction (WERR) for intra-domain LibriSpeech data, as well as 21.4% and 24.6% relative WERRs on cross-domain TED-LIUM 2 and AESRC2020 data with an adapted prediction network.

📄 PDF Abstract BibTeX arXiv:2311.11353

Code (0)

등록된 구현이 없습니다.

Tasks

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Domain AdaptationLanguage ModellingPredictionspeech-recognitionSpeech Recognition

Similar Papers 제목 키워드 기반

Label-Synchronous Neural Transducer for E2E Simultaneous Speech Translation

2024-06-06 · Keqi Deng, Philip C. Woodland

While the neural transducer is popular for online speech recognition, simultaneous speech translation (SST) requires both streaming and re-ordering capabilities. This paper presents the LS-Transducer-SST, a label-synchro…

es-enspeech-recognitionSpeech RecognitionTranslation

Label-Synchronous Neural Transducer for End-to-End ASR

2023-07-06 · Keqi Deng, Philip C. Woodland

Neural transducers provide a natural way of streaming ASR. However, they augment output sequences with blank tokens which leads to challenges for domain adaptation using text data. This paper proposes a label-synchronous…

Domain AdaptationPrediction

Equivalence of Segmental and Neural Transducer Modeling: A Proof of Concept

2021-04-13 · Wei Zhou, Albert Zeyer, André Merboldt, Ralf Schlüter 외

With the advent of direct models in automatic speech recognition (ASR), the formerly prevalent frame-wise acoustic modeling based on hidden Markov models (HMM) diversified into a number of modeling architectures like enc…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Decoderspeech-recognition+1

TRADE: Transducer-Augmented Decoder for Speech LLM

2026-06-07 · Yun Tang, Shanil Puri, Shinji Watanabe, Subhabrata Mukherjee arxiv

Speech Large Language Models (Speech LLMs) lack a principled mechanism for streaming inference: their label-synchronous generation has no acoustic-frame alignment, making real-time decoding and end-of-utterance detection…

Activity Detection

Integration of Frame- and Label-synchronous Beam Search for Streaming Encoder-decoder Speech Recognition

2023-07-24 · Emiru Tsunoo, Hayato Futami, Yosuke Kashiwagi, Siddhant Arora 외

Although frame-based models, such as CTC and transducers, have an affinity for streaming automatic speech recognition, their decoding uses no future knowledge, which could lead to incorrect pruning. Conversely, label-bas…

Automatic Speech RecognitionDecoderspeech-recognitionSpeech Recognition