paper-with-me

Papers

Improving noisy student training for low-resource languages in End-to-End ASR using CycleGAN and inter-domain losses

2024-07-26 · Chia-Yu Li, Ngoc Thang Vu

Training a semi-supervised end-to-end speech recognition system using noisy student training has significantly improved performance. However, this approach requires a substantial amount of paired speech-text and unlabeled speech, which is costly for low-resource languages. Therefore, this paper considers a more extreme case of semi-supervised end-to-end automatic speech recognition where there are limited paired speech-text, unlabeled speech (less than five hours), and abundant external text. Firstly, we observe improved performance by training the model using our previous work on semi-supervised learning "CycleGAN and inter-domain losses" solely with external text. Secondly, we enhance "CycleGAN and inter-domain losses" by incorporating automatic hyperparameter tuning, calling it "enhanced CycleGAN inter-domain losses." Thirdly, we integrate it into the noisy student training approach pipeline for low-resource scenarios. Our experimental results, conducted on six non-English languages from Voxforge and Common Voice, show a 20% word error rate reduction compared to the baseline teacher model and a 10% word error rate reduction compared to the baseline best student model, highlighting the significant improvements achieved through our proposed method.

📄 PDF Abstract BibTeX arXiv:2407.21061

Code (0)

등록된 구현이 없습니다.

Tasks

Automatic Speech Recognitionspeech-recognitionSpeech Recognition

Methods 이 논문이 사용한 방법론

HuMan(Expedia)||How do I get a human at Expedia? How do I get a human at Expedia? How Do I Get a Human at Expedia? – Call ☎️ +1-(888) 829 (0881) or +1-805-330-4056 or +1-805-330-4056 for Real-Time Help & Exclusive…
ReLU How Do I Communicate to Expedia? How Do I Communicate to Expedia? – Call ☎️ +1-(888) 829 (0881) or +1-805-330-4056 or +1-805-330-4056 for Live Support & Special Travel…
Instance Normalization Instance Normalization (also known as contrast normalization) is a normalization layer where: $$ y_{tijk} = \frac{x_{tijk} - \mu_{ti}}{\sqrt{\sigma_{ti}^2 +…
Batch Normalization 설명 없음
PatchGAN 설명 없음
Residual Connection 설명 없음
RandAugment 설명 없음
Sigmoid Activation 설명 없음

Similar Papers 제목 키워드 기반

Speech Enhancement Based on Cyclegan with Noise-informed Training

2021-10-19 · Wen-Yuan Ting, Syu-Siang Wang, Hsin-Li Chang, Borching Su 외

Cycle-consistent generative adversarial networks (CycleGAN) were successfully applied to speech enhancement (SE) tasks with unpaired noisy-clean training data. The CycleGAN SE system adopted two generators and two discri…

Speech Enhancement

You Can Have Your Data and Balance It Too: Towards Balanced and Efficient Multilingual Models

2022-10-13 · Tomasz Limisiewicz, Dan Malkin, Gabriel Stanovsky

Multilingual models have been widely used for cross-lingual transfer to low-resource languages. However, the performance on these languages is hindered by their underrepresentation in the pretraining data. To alleviate t…

Cross-Lingual TransferKnowledge Distillation

Adversarial and Score-Based CT Denoising: CycleGAN vs Noise2Score

2025-11-06 · Abu Hanif Muhammad Syarubany arxiv

We study CT image denoising in the unpaired and self-supervised regimes by evaluating two strong, training-data-efficient paradigms: a CycleGAN-based residual translator and a Noise2Score (N2S) score-matching denoiser. U…

Image Denoising

Improving Speech Recognition on Noisy Speech via Speech Enhancement with Multi-Discriminators CycleGAN

2021-12-12 · Chia-Yu Li, Ngoc Thang Vu

This paper presents our latest investigations on improving automatic speech recognition for noisy speech via speech enhancement. We propose a novel method named Multi-discriminators CycleGAN to reduce noise of input spee…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Speech Enhancementspeech-recognition+1

Bitext Mining Using Distilled Sentence Representations for Low-Resource Languages

2022-05-25 · Kevin Heffernan, Onur Çelebi, Holger Schwenk

Scaling multilingual representation learning beyond the hundred most frequent languages is challenging, in particular to cover the long tail of low-resource languages. A promising approach has been to train one-for-all m…

Cross-Lingual TransferNMTRepresentation LearningSentence