Improving End-to-end Speech Translation by Leveraging Auxiliary Speech and Text Data
We present a method for introducing a text encoder into pre-trained end-to-end speech translation systems. It enhances the ability of adapting one modality (i.e., source-language speech) to another (i.e., source-language text). Thus, the speech translation model can learn from both unlabeled and labeled data, especially when the source-language text data is abundant. Beyond this, we present a denoising method to build a robust text encoder that can deal with both normal and noisy text data. Our system sets new state-of-the-arts on the MuST-C En-De, En-Fr, and LibriSpeech En-Fr tasks.
Code (0)
등록된 구현이 없습니다.
Tasks
de-enDenoisingTranslationSimilar Papers 제목 키워드 기반
Improving End-to-end Speech Translation by Leveraging Auxiliary Speech and Text Data
We present a method for introducing a text encoder into pre-training end-to-end speech translation systems. It enhances the ability of adapting one modality (i.e., source-language speech) to another (i.e., source-languag…
de-enDenoisingTranslationDirect Speech-to-speech Translation without Textual Annotation using Bottleneck Features
Speech-to-speech translation directly translates a speech utterance to another between different languages, and has great potential in tasks such as simultaneous interpretation. State-of-art models usually contains an au…
Speech-to-Speech TranslationTranslationTextless Direct Speech-to-Speech Translation with Discrete Speech Representation
Research on speech-to-speech translation (S2ST) has progressed rapidly in recent years. Many end-to-end systems have been proposed and show advantages over conventional cascade systems, which are often composed of recogn…
Speech-to-Speech TranslationTranslationMultilingual Speech-to-Speech Translation into Multiple Target Languages
Speech-to-speech translation (S2ST) enables spoken communication between people talking in different languages. Despite a few studies on multilingual S2ST, their focus is the multilinguality on the source side, i.e., the…
Language IdentificationSpeech-to-Speech TranslationTranslationUnified Speech-Text Pre-training for Speech Translation and Recognition
We describe a method to jointly pre-train speech and text in an encoder-decoder modeling framework for speech translation and recognition. The proposed method incorporates four self-supervised and supervised subtasks for…
Decoderspeech-recognitionSpeech RecognitionTranslation