Improving Cross-Lingual Transfer Learning for End-to-End Speech Recognition with Speech Translation
Transfer learning from high-resource languages is known to be an efficient way to improve end-to-end automatic speech recognition (ASR) for low-resource languages. Pre-trained or jointly trained encoder-decoder models, however, do not share the language modeling (decoder) for the same language, which is likely to be inefficient for distant target languages. We introduce speech-to-text translation (ST) as an auxiliary task to incorporate additional knowledge of the target language and enable transferring from that target language. Specifically, we first translate high-resource ASR transcripts into a target low-resource language, with which a ST model is trained. Both ST and target ASR share the same attention-based encoder-decoder architecture and vocabulary. The former task then provides a fully pre-trained model for the latter, bringing up to 24.6% word error rate (WER) reduction to the baseline (direct transfer from high-resource ASR). We show that training ST with human translations is not necessary. ST trained with machine translation (MT) pseudo-labels brings consistent gains. It can even outperform those using human labels when transferred to target ASR by leveraging only 500K MT examples. Even with pseudo-labels from low-resource MT (200K examples), ST-enhanced transfer brings up to 8.9% WER reduction to direct transfer.
Code (0)
등록된 구현이 없습니다.
Tasks
Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Cross-Lingual TransferDecoderLanguage ModelingLanguage ModellingMachine Translationspeech-recognitionSpeech RecognitionSpeech-to-TextSpeech-to-Text TranslationTransfer LearningTranslationSimilar Papers 제목 키워드 기반
A Survey of Multilingual Models for Automatic Speech Recognition
Although Automatic Speech Recognition (ASR) systems have achieved human-like performance for a few languages, the majority of the world's languages do not have usable systems due to the lack of large speech datasets to t…
Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Cross-Lingual TransferSelf-Supervised Learning+4Cross-lingual Transfer for Speech Processing using Acoustic Language Similarity
Speech processing systems currently do not support the vast majority of languages, in part due to the lack of data in low-resource languages. Cross-lingual transfer offers a compelling way to help bridge this digital div…
Cross-Lingual Transferspeech-recognitionSpeech RecognitionSpeech SynthesisCross-Lingual Cross-Age Group Adaptation for Low-Resource Elderly Speech Emotion Recognition
Speech emotion recognition plays a crucial role in human-computer interactions. However, most speech emotion recognition research is biased toward English-speaking adults, which hinders its applicability to other demogra…
Data AugmentationEmotion RecognitionSpeech Emotion RecognitionMultilingual Auxiliary Tasks Training: Bridging the Gap between Languages for Zero-Shot Transfer of Hate Speech Detection Models
Zero-shot cross-lingual transfer learning has been shown to be highly challenging for tasks involving a lot of linguistic specificities or when a cultural gap is present between languages, such as in hate speech detectio…
Cross-Lingual TransferHate Speech Detectionnamed-entity-recognitionNamed Entity Recognition+4When More is not Necessary Better: Multilingual Auxiliary Tasks for Zero-Shot Cross-Lingual Transfer of Hate Speech Detection Models
Zero-shot cross-lingual transfer learning has been shown to be highly challenging for tasks involving a lot of linguistic specificities or when a cultural gap is present between languages, such as in hate speech detectio…
Cross-Lingual TransferHate Speech DetectionLanguage ModelingLanguage Modelling+6