Simple and Effective Unsupervised Speech Translation
The amount of labeled data to train models for speech tasks is limited for most languages, however, the data scarcity is exacerbated for speech translation which requires labeled data covering two different languages. To address this issue, we study a simple and effective approach to build speech translation systems without labeled data by leveraging recent advances in unsupervised speech recognition, machine translation and speech synthesis, either in a pipeline approach, or to generate pseudo-labels for training end-to-end speech translation models. Furthermore, we present an unsupervised domain adaptation technique for pre-trained speech models which improves the performance of downstream unsupervised speech recognition, especially for low-resource settings. Experiments show that unsupervised speech-to-text translation outperforms the previous unsupervised state of the art by 3.2 BLEU on the Libri-Trans benchmark, on CoVoST 2, our best systems outperform the best supervised end-to-end models (without pre-training) from only two years ago by an average of 5.0 BLEU over five X-En directions. We also report competitive results on MuST-C and CVSS benchmarks.
Code (0)
등록된 구현이 없습니다.
Tasks
Domain AdaptationMachine Translationspeech-recognitionSpeech RecognitionSpeech SynthesisSpeech-to-TextSpeech-to-Text TranslationTranslationUnsupervised Domain AdaptationUnsupervised Speech RecognitionSimilar Papers 제목 키워드 기반
Towards speech-to-text translation without speech recognition
We explore the problem of translating speech to text in low-resource scenarios where neither automatic speech recognition (ASR) nor machine translation (MT) are available, but we have training data in the form of audio p…
Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Machine Translationspeech-recognition+4Simple and Effective Unsupervised Speech Synthesis
We introduce the first unsupervised speech synthesis system based on a simple, yet effective recipe. The framework leverages recent work in unsupervised speech recognition as well as existing neural-based speech synthesi…
speech-recognitionSpeech RecognitionSpeech SynthesisUnsupervised Speech RecognitionImproving Cascaded Unsupervised Speech Translation with Denoising Back-translation
Most of the speech translation models heavily rely on parallel data, which is hard to collect especially for low-resource languages. To tackle this issue, we propose to build a cascaded speech translation system without …
DenoisingMachine TranslationTranslationTranslatotron 3: Speech to Speech Translation with Monolingual Data
This paper presents Translatotron 3, a novel approach to unsupervised direct speech-to-speech translation from monolingual speech-text datasets by combining masked autoencoder, unsupervised embedding mapping, and back-tr…
Speech-to-Speech TranslationTranslationLeveraging unsupervised and weakly-supervised data to improve direct speech-to-speech translation
End-to-end speech-to-speech translation (S2ST) without relying on intermediate text representations is a rapidly emerging frontier of research. Recent works have demonstrated that the performance of such direct S2ST syst…
Representation LearningSpeech Representation LearningSpeech-to-Speech TranslationTranslation