PolyVoice: Language Models for Speech to Speech Translation
We propose PolyVoice, a language model-based framework for speech-to-speech translation (S2ST) system. Our framework consists of two language models: a translation language model and a speech synthesis language model. We use discretized speech units, which are generated in a fully unsupervised way, and thus our framework can be used for unwritten languages. For the speech synthesis part, we adopt the existing VALL-E X approach and build a unit-based audio language model. This grants our framework the ability to preserve the voice characteristics and the speaking style of the original speech. We examine our system on Chinese $\rightarrow$ English and English $\rightarrow$ Spanish pairs. Experimental results show that our system can generate speech with high translation quality and audio quality. Speech samples are available at https://speechtranslation.github.io/polyvoice.
Code (0)
등록된 구현이 없습니다.
Tasks
Language ModelingLanguage ModellingSpeech SynthesisSpeech-to-Speech TranslationTranslationSimilar Papers 제목 키워드 기반
Assessing Evaluation Metrics for Speech-to-Speech Translation
Speech-to-speech translation combines machine translation with speech synthesis, introducing evaluation challenges not present in either task alone. How to automatically evaluate speech-to-speech translation is an open q…
Machine TranslationOpen-Ended Question AnsweringSpeech SynthesisSpeech-to-Speech Translation+1TranSentence: Speech-to-speech Translation via Language-agnostic Sentence-level Speech Encoding without Language-parallel Data
Although there has been significant advancement in the field of speech-to-speech translation, conventional models still require language-parallel speech data between the source and target languages for training. In this …
SentenceSpeech-to-Speech TranslationTranslationSeamlessM4T: Massively Multilingual & Multimodal Machine Translation
What does it take to create the Babel Fish, a tool that can help individuals translate speech between any two languages? While recent breakthroughs in text-based models have pushed machine translation coverage beyond 200…
Automatic Speech RecognitionMachine TranslationSpeech-to-Speech TranslationSpeech-to-Text+5UWSpeech: Speech to Speech Translation for Unwritten Languages
Existing speech to speech translation systems heavily rely on the text of target language: they usually translate source language either to target text and then synthesize target speech from text, or directly to target s…
speech-recognitionSpeech RecognitionSpeech-to-Speech TranslationTranslationCode-Switching without Switching: Language Agnostic End-to-End Speech Translation
We propose a) a Language Agnostic end-to-end Speech Translation model (LAST), and b) a data augmentation strategy to increase code-switching (CS) performance. With increasing globalization, multiple languages are increas…
Data Augmentationspeech-recognitionSpeech RecognitionTranslation