paper-with-me

홈 › Papers

Pushing the Limits of Semi-Supervised Learning for Automatic Speech Recognition

2020-10-20 · Yu Zhang, James Qin, Daniel S. Park, Wei Han, Chung-Cheng Chiu, Ruoming Pang, Quoc V. Le, Yonghui Wu

We employ a combination of recent developments in semi-supervised learning for automatic speech recognition to obtain state-of-the-art results on LibriSpeech utilizing the unlabeled audio of the Libri-Light dataset. More precisely, we carry out noisy student training with SpecAugment using giant Conformer models pre-trained using wav2vec 2.0 pre-training. By doing so, we are able to achieve word-error-rates (WERs) 1.4%/2.6% on the LibriSpeech test/test-other sets against the current state-of-the-art WERs 1.7%/3.3%.

📄 PDF Abstract BibTeX arXiv:2010.10504

Code (1)

tuanio/noisy-student-training-asr pytorch

Tasks

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)speech-recognitionSpeech Recognition

Methods 이 논문이 사용한 방법론

Dropout Dropout is a regularization technique for neural networks that drops a unit (along with connections) at training time with a specified probability $p$ (a common value is…
Stochastic Depth Stochastic Depth aims to shrink the depth of a network during training, while keeping it unchanged during testing. This is achieved by randomly dropping entire…
RandAugment 설명 없음
Noisy Student 설명 없음

Similar Papers 제목 키워드 기반

Pushing the Limits of Non-Autoregressive Speech Recognition

2021-04-07 · Edwin G. Ng, Chung-Cheng Chiu, Yu Zhang, William Chan

We combine recent advancements in end-to-end speech recognition to non-autoregressive automatic speech recognition. We push the limits of non-autoregressive state-of-the-art results for multiple datasets: LibriSpeech, Fi…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Language ModelingLanguage Modelling+2

Semi-supervised Acoustic Modelling for Five-lingual Code-switched ASR using Automatically-segmented Soap Opera Speech

2020-05-01 · LREC 2020 5 · Nick Wilkinson, Astik Biswas, Emre Yilmaz, Febe De Wet 외

This paper considers the impact of automatic segmentation on the fully-automatic, semi-supervised training of automatic speech recog-nition (ASR) systems for five-lingual code-switched (CS) speech. Four automatic segment…

Acoustic ModellingAction DetectionActivity DetectionSegmentation+2

Improving End-to-End Bangla Speech Recognition with Semi-supervised Training

2020-11-01 · Findings of the Association for Computational Linguistics 2020 · Nafis Sadeq, Nafis Tahmid Chowdhury, Farhan Tanvir Utshaw, Shafayat Ahmed 외

Automatic speech recognition systems usually require large annotated speech corpus for training. The manual annotation of a large corpus is very difficult. It can be very helpful to use unsupervised and semi-supervised l…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)speech-recognitionSpeech Recognition

Pushing the Limits of Unsupervised Unit Discovery for SSL Speech Representation

2023-06-15 · Ziyang Ma, Zhisheng Zheng, Guanrou Yang, Yu Wang 외

The excellent generalization ability of self-supervised learning (SSL) for speech foundation models has garnered significant attention. HuBERT is a successful example that utilizes offline clustering to convert speech fe…

Automatic Speech RecognitionClusteringLanguage ModelingLanguage Modelling+3

Semi-supervised acoustic modelling for five-lingual code-switched ASR using automatically-segmented soap opera speech

2020-04-08 · N. Wilkinson, A. Biswas, E. Yılmaz, F. de Wet 외

This paper considers the impact of automatic segmentation on the fully-automatic, semi-supervised training of automatic speech recognition (ASR) systems for five-lingual code-switched (CS) speech. Four automatic segmenta…

Acoustic ModellingAction DetectionActivity DetectionAutomatic Speech Recognition+6