Unified Speech-Text Pre-training for Speech Translation and Recognition
In this work, we describe a method to jointly pre-train speech and text in an encoder-decoder modeling framework for speech translation and recognition. The proposed method utilizes multi-task learning to integrate four self-supervised and supervised subtasks for cross modality learning. A self-supervised speech subtask, which leverages unlabelled speech data, and a (self-)supervised text to text subtask, which makes use of abundant text training data, take up the majority of the pre-training time. Two auxiliary supervised speech tasks are included to unify speech and text modeling space. Detailed analysis reveals learning interference among subtasks. In order to alleviate the subtask interference, two pre-training configurations are proposed for speech translation and speech recognition respectively. Our experiments show the proposed method can effectively fuse speech and text information into one model. It achieves between 1.7 and 2.3 BLEU improvement above the state of the art on the MuST-C speech translation dataset and comparable WERs to wav2vec 2.0 on the Librispeech speech recognition task.
Code (0)
등록된 구현이 없습니다.
Tasks
DecoderMulti-Task Learningspeech-recognitionSpeech RecognitionTranslationSimilar Papers 제목 키워드 기반
Fused Acoustic and Text Encoding for Multimodal Bilingual Pretraining and Speech Translation
Recently, representation learning for text and speech has successfully improved many language related tasks. However, all existing methods suffer from two limitations: (a) they only learn from one input modality, while a…
Language ModelingLanguage ModellingMachine TranslationRepresentation Learning+3Unified Speech-Text Pre-training for Speech Translation and Recognition
We describe a method to jointly pre-train speech and text in an encoder-decoder modeling framework for speech translation and recognition. The proposed method incorporates four self-supervised and supervised subtasks for…
Decoderspeech-recognitionSpeech RecognitionTranslationSeamlessM4T: Massively Multilingual & Multimodal Machine Translation
What does it take to create the Babel Fish, a tool that can help individuals translate speech between any two languages? While recent breakthroughs in text-based models have pushed machine translation coverage beyond 200…
Automatic Speech RecognitionMachine TranslationSpeech-to-Speech TranslationSpeech-to-Text+5Multilingual Speech Translation with Unified Transformer: Huawei Noah’s Ark Lab at IWSLT 2021
This paper describes the system submitted to the IWSLT 2021 Multilingual Speech Translation (MultiST) task from Huawei Noah’s Ark Lab. We use a unified transformer architecture for our MultiST model, so that the data fro…
Data AugmentationDecoderMachine TranslationMulti-Task Learning+3Multilingual Speech Translation with Unified Transformer: Huawei Noah's Ark Lab at IWSLT 2021
This paper describes the system submitted to the IWSLT 2021 Multilingual Speech Translation (MultiST) task from Huawei Noah's Ark Lab. We use a unified transformer architecture for our MultiST model, so that the data fro…
Data AugmentationDecoderMachine TranslationMulti-Task Learning+3