Joint Pre-Training with Speech and Bilingual Text for Direct Speech to Speech Translation
Direct speech-to-speech translation (S2ST) is an attractive research topic with many advantages compared to cascaded S2ST. However, direct S2ST suffers from the data scarcity problem because the corpora from speech of the source language to speech of the target language are very rare. To address this issue, we propose in this paper a Speech2S model, which is jointly pre-trained with unpaired speech and bilingual text data for direct speech-to-speech translation tasks. By effectively leveraging the paired text data, Speech2S is capable of modeling the cross-lingual speech conversion from source to target language. We verify the performance of the proposed Speech2S on Europarl-ST and VoxPopuli datasets. Experimental results demonstrate that Speech2S gets an improvement of about 5 BLEU scores compared to encoder-only pre-training models, and achieves a competitive or even better performance than existing state-of-the-art models1.
Code (1)
Tasks
Speech-to-Speech TranslationTranslationSimilar Papers 제목 키워드 기반
Textless Direct Speech-to-Speech Translation with Discrete Speech Representation
Research on speech-to-speech translation (S2ST) has progressed rapidly in recent years. Many end-to-end systems have been proposed and show advantages over conventional cascade systems, which are often composed of recogn…
Speech-to-Speech TranslationTranslationEnd-to-end Speech Translation with Spoken-to-Written Style Conversion
End-to-end speech translation (ST), which translates speech in source language directly into text in target language by a single model, has attracted a great deal of attention in recent years. Compared to the cascade ST,…
DecoderMachine TranslationTranslationJoint Modeling of Code-Switched and Monolingual ASR via Conditional Factorization
Conversational bilingual speech encompasses three types of utterances: two purely monolingual types and one intra-sententially code-switched type. In this work, we propose a general framework to jointly model the likelih…
speech-recognitionSpeech RecognitionFused Acoustic and Text Encoding for Multimodal Bilingual Pretraining and Speech Translation
Recently, representation learning for text and speech has successfully improved many language related tasks. However, all existing methods suffer from two limitations: (a) they only learn from one input modality, while a…
Language ModelingLanguage ModellingMachine TranslationRepresentation Learning+3Improving low-resource ASR using bilingual fine-tuning with language identification: a cross-linguistic evaluation
This study explores how bilingual fine-tuning affects automatic speech recognition (ASR) in low-resource languages. We evaluate this method across nine linguistically and geographically diverse language pairs, covering a…
Language IdentificationSpeech Recognition