paper-with-me

Papers

Unified Semi-Supervised Pipeline for Automatic Speech Recognition

2025-06-09 · Nune Tadevosyan, Nikolay Karpov, Andrei Andrusenko, Vitaly Lavrukhin, Ante Jukic

Automatic Speech Recognition has been a longstanding research area, with substantial efforts dedicated to integrating semi-supervised learning due to the scarcity of labeled datasets. However, most prior work has focused on improving learning algorithms using existing datasets, without providing a complete public framework for large-scale semi-supervised training across new datasets or languages. In this work, we introduce a fully open-source semi-supervised training framework encompassing the entire pipeline: from unlabeled data collection to pseudo-labeling and model training. Our approach enables scalable dataset creation for any language using publicly available speech data under Creative Commons licenses. We also propose a novel pseudo-labeling algorithm, TopIPL, and evaluate it in both low-resource (Portuguese, Armenian) and high-resource (Spanish) settings. Notably, TopIPL achieves relative WER improvements of 18-40% for Portuguese, 5-16% for Armenian, and 2-8% for Spanish.

📄 PDF Abstract BibTeX arXiv:2506.07659

Code (0)

등록된 구현이 없습니다.

Tasks

Automatic Speech Recognitionspeech-recognitionSpeech Recognition

Similar Papers 제목 키워드 기반

Semi-supervised acoustic model training for five-lingual code-switched ASR

2019-06-20 · Astik Biswas, Emre Yilmaz, Febe De Wet, Ewald van der Westhuizen 외

This paper presents recent progress in the acoustic modelling of under-resourced code-switched (CS) speech in multiple South African languages. We consider two approaches. The first constructs separate bilingual acoustic…

Acoustic ModellingLanguage ModelingLanguage Modelling

Semi-supervised Development of ASR Systems for Multilingual Code-switched Speech in Under-resourced Languages

2020-03-06 · LREC 2020 5 · Astik Biswas, Emre Yilmaz, Febe De Wet, Ewald van der Westhuizen 외

This paper reports on the semi-supervised development of acoustic and language models for under-resourced, code-switched speech in five South African languages. Two approaches are considered. The first constructs four se…

Better Pseudo-labeling with Multi-ASR Fusion and Error Correction by SpeechLLM

2025-06-05 · Jeena Prakash, Blessingh Kumar, Kadri Hacioglu, Bidisha Sharma 외

Automatic speech recognition (ASR) models rely on high-quality transcribed data for effective training. Generating pseudo-labels for large unlabeled audio datasets often relies on complex pipelines that combine multiple …

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)speech-recognitionSpeech Recognition

Improving noisy student training for low-resource languages in End-to-End ASR using CycleGAN and inter-domain losses

2024-07-26 · Chia-Yu Li, Ngoc Thang Vu

Training a semi-supervised end-to-end speech recognition system using noisy student training has significantly improved performance. However, this approach requires a substantial amount of paired speech-text and unlabele…

Automatic Speech Recognitionspeech-recognitionSpeech Recognition

StreamHear: Domain-Adapted Pseudo-Labeling for Semi-Supervised Streaming Speech Recognition

2026-08-13 · Zefang Liu, Chenyang Zhu, Sangwoo Cho, Xujun Peng 외 arxiv

Streaming automatic speech recognition (ASR) underperforms on domain-shifted target audio, where labeled in-domain data is costly to prepare while unlabeled audio is abundant. We present StreamHear, a semi-supervised pip…

Speech Recognition