Translatotron 2: High-quality direct speech-to-speech translation with voice preservation
We present Translatotron 2, a neural direct speech-to-speech translation model that can be trained end-to-end. Translatotron 2 consists of a speech encoder, a linguistic decoder, an acoustic synthesizer, and a single attention module that connects them together. Experimental results on three datasets consistently show that Translatotron 2 outperforms the original Translatotron by a large margin on both translation quality (up to +15.5 BLEU) and speech generation quality, and approaches the same of cascade systems. In addition, we propose a simple method for preserving speakers' voices from the source speech to the translation speech in a different language. Unlike existing approaches, the proposed method is able to preserve each speaker's voice on speaker turns without requiring for speaker segmentation. Furthermore, compared to existing approaches, it better preserves speaker's privacy and mitigates potential misuse of voice cloning for creating spoofing audio artifacts.
Code (0)
등록된 구현이 없습니다.
Tasks
Data AugmentationDecoderSpeech-to-Speech TranslationTranslationVoice CloningSimilar Papers 제목 키워드 기반
Speech to Speech Translation with Translatotron: A State of the Art Review
A cascade-based speech-to-speech translation has been considered a benchmark for a very long time, but it is plagued by many issues, like the time taken to translate a speech from one language to another and compound err…
speech-recognitionSpeech RecognitionSpeech-to-Speech TranslationSpeech-to-Text+5Textless Direct Speech-to-Speech Translation with Discrete Speech Representation
Research on speech-to-speech translation (S2ST) has progressed rapidly in recent years. Many end-to-end systems have been proposed and show advantages over conventional cascade systems, which are often composed of recogn…
Speech-to-Speech TranslationTranslationTranslatotron 3: Speech to Speech Translation with Monolingual Data
This paper presents Translatotron 3, a novel approach to unsupervised direct speech-to-speech translation from monolingual speech-text datasets by combining masked autoencoder, unsupervised embedding mapping, and back-tr…
Speech-to-Speech TranslationTranslationSimulTron: On-Device Simultaneous Speech to Speech Translation
Simultaneous speech-to-speech translation (S2ST) holds the promise of breaking down communication barriers and enabling fluid conversations across languages. However, achieving accurate, real-time translation through mob…
Simultaneous Speech-to-Speech TranslationSpeech-to-Speech TranslationTranslationCan We Achieve High-quality Direct Speech-to-Speech Translation without Parallel Speech Data?
Recently proposed two-pass direct speech-to-speech translation (S2ST) models decompose the task into speech-to-text translation (S2TT) and text-to-speech (TTS) within an end-to-end model, yielding promising results. Howe…
Contrastive LearningSpeech SynthesisSpeech-to-Speech TranslationSpeech-to-Text+5