paper-with-me

Papers

WST: Weakly Supervised Transducer for Automatic Speech Recognition

2025-11-06 · Dongji Gao, Chenda Liao, Changliang Liu, Matthew Wiesner, Leibny Paola Garcia, Daniel Povey, Sanjeev Khudanpur, Jian Wu arxiv

The Recurrent Neural Network-Transducer (RNN-T) is widely adopted in end-to-end (E2E) automatic speech recognition (ASR) tasks but depends heavily on large-scale, high-quality annotated data, which are often costly and difficult to obtain. To mitigate this reliance, we propose a Weakly Supervised Transducer (WST), which integrates a flexible training graph designed to robustly handle errors in the transcripts without requiring additional confidence estimation or auxiliary pre-trained models. Empirical evaluations on synthetic and industrial datasets reveal that WST effectively maintains performance even with transcription error rates of up to 70%, consistently outperforming existing Connectionist Temporal Classification (CTC)-based weakly supervised approaches, such as Bypass Temporal Classification (BTC) and Omni-Temporal Classification (OTC). These results demonstrate the practical utility and robustness of WST in realistic ASR settings. The implementation will be publicly available.

📄 PDF Abstract BibTeX arXiv:2511.04035

Code (0)

등록된 구현이 없습니다.

Tasks

Speech Recognition

Similar Papers 제목 키워드 기반

Bypass Temporal Classification: Weakly Supervised Automatic Speech Recognition with Imperfect Transcripts

2023-06-01 · Dongji Gao, Matthew Wiesner, Hainan Xu, Leibny Paola Garcia 외

This paper presents a novel algorithm for building an automatic speech recognition (ASR) model with imperfect training data. Imperfectly transcribed speech is a prevalent issue in human-annotated speech corpora, which de…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)speech-recognitionSpeech Recognition

Learning from Flawed Data: Weakly Supervised Automatic Speech Recognition

2023-09-26 · Dongji Gao, Hainan Xu, Desh Raj, Leibny Paola Garcia Perera 외

Training automatic speech recognition (ASR) systems requires large amounts of well-curated paired data. However, human annotators usually perform "non-verbatim" transcription, which can result in poorly trained models. I…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)speech-recognitionSpeech Recognition

A comparative analysis between Conformer-Transducer, Whisper, and wav2vec2 for improving the child speech recognition

2023-11-07 · Andrei Barcovschi, Rishabh Jain, Peter Corcoran

Automatic Speech Recognition (ASR) systems have progressed significantly in their performance on adult speech data; however, transcribing child speech remains challenging due to the acoustic differences in the characteri…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)speech-recognitionSpeech Recognition

Open Source State-Of-the-Art Solution for Romanian Speech Recognition

2025-11-05 · Gabriel Pirlogeanu, Alexandru-Lucian Georgescu, Horia Cucu arxiv

In this work, we present a new state-of-the-art Romanian Automatic Speech Recognition (ASR) system based on NVIDIA's FastConformer architecture--explored here for the first time in the context of Romanian. We train our m…

Speech Recognition

A Weakly-Supervised Streaming Multilingual Speech Model with Truly Zero-Shot Capability

2022-11-04 · Jian Xue, Peidong Wang, Jinyu Li, Eric Sun

In this paper, we introduce our work of building a Streaming Multilingual Speech Model (SM2), which can transcribe or translate multiple spoken languages into texts of the target language. The backbone of SM2 is Transfor…

Machine Translationspeech-recognitionSpeech RecognitionTranslation