Self-Training for End-to-End Speech Translation
One of the main challenges for end-to-end speech translation is data scarcity. We leverage pseudo-labels generated from unlabeled audio by a cascade and an end-to-end speech translation model. This provides 8.3 and 5.7 BLEU gains over a strong semi-supervised baseline on the MuST-C English-French and English-German datasets, reaching state-of-the art performance. The effect of the quality of the pseudo-labels is investigated. Our approach is shown to be more effective than simply pre-training the encoder on the speech recognition task. Finally, we demonstrate the effectiveness of self-training by directly generating pseudo-labels with an end-to-end model instead of a cascade model.
Code (0)
등록된 구현이 없습니다.
Tasks
speech-recognitionSpeech RecognitionTranslationSimilar Papers 제목 키워드 기반
Enhanced Direct Speech-to-Speech Translation Using Self-supervised Pre-training and Data Augmentation
Direct speech-to-speech translation (S2ST) models suffer from data scarcity issues as there exists little parallel S2ST data, compared to the amount of data available for conventional cascaded systems that consist of aut…
Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Data AugmentationDecoder+9Unified Speech-Text Pre-training for Speech Translation and Recognition
We describe a method to jointly pre-train speech and text in an encoder-decoder modeling framework for speech translation and recognition. The proposed method incorporates four self-supervised and supervised subtasks for…
Decoderspeech-recognitionSpeech RecognitionTranslationUnified Speech-Text Pre-training for Speech Translation and Recognition
In this work, we describe a method to jointly pre-train speech and text in an encoder-decoder modeling framework for speech translation and recognition. The proposed method utilizes multi-task learning to integrate fou…
DecoderMulti-Task Learningspeech-recognitionSpeech Recognition+1Self-Supervised Representations Improve End-to-End Speech Translation
End-to-end speech-to-text translation can provide a simpler and smaller system but is facing the challenge of data scarcity. Pre-training methods can leverage unlabeled data and have been shown to be effective on data-sc…
Cross-Lingual Transferspeech-recognitionSpeech RecognitionSpeech-to-Text+2Textless Acoustic Model with Self-Supervised Distillation for Noise-Robust Expressive Speech-to-Speech Translation
In this paper, we propose a textless acoustic model with a self-supervised distillation strategy for noise-robust expressive speech-to-speech translation (S2ST). Recently proposed expressive S2ST systems have achieved im…
Speech-to-Speech TranslationTranslation