Diffusion Synthesizer for Efficient Multilingual Speech to Speech Translation
We introduce DiffuseST, a low-latency, direct speech-to-speech translation system capable of preserving the input speaker's voice zero-shot while translating from multiple source languages into English. We experiment with the synthesizer component of the architecture, comparing a Tacotron-based synthesizer to a novel diffusion-based synthesizer. We find the diffusion-based synthesizer to improve MOS and PESQ audio quality metrics by 23\% each and speaker similarity by 5\% while maintaining comparable BLEU scores. Despite having more than double the parameter count, the diffusion synthesizer has lower latency, allowing the entire model to run more than 5$\times$ faster than real-time.
Code (0)
등록된 구현이 없습니다.
Tasks
Speech-to-Speech TranslationTranslationMethods 이 논문이 사용한 방법론
Similar Papers 제목 키워드 기반
DiffSSD: A Diffusion-Based Dataset For Speech Forensics
Diffusion-based speech generators are ubiquitous. These methods can generate very high quality synthetic speech and several recent incidents report their malicious use. To counter such misuse, synthetic speech detectors …
TransFace: Unit-Based Audio-Visual Speech Synthesizer for Talking Head Translation
Direct speech-to-speech translation achieves high-quality results through the introduction of discrete units obtained from self-supervised learning. This approach circumvents delays and cascading errors associated with m…
es-enfr-enSelf-Supervised LearningSpeech-to-Speech Translation+1SpeechMatrix: A Large-Scale Mined Corpus of Multilingual Speech-to-Speech Translations
We present SpeechMatrix, a large-scale multilingual corpus of speech-to-speech translations mined from real speech of European Parliament recordings. It contains speech alignments in 136 language pairs with a total of 41…
Mixture-of-ExpertsSpeech-to-Speech TranslationTranslationCosyVoice: A Scalable Multilingual Zero-shot Text-to-speech Synthesizer based on Supervised Semantic Tokens
Recent years have witnessed a trend that large language model (LLM) based text-to-speech (TTS) emerges into the mainstream due to their high naturalness and zero-shot capacity. In this paradigm, speech signals are discre…
Language ModellingLarge Language ModelQuantizationspeech-recognition+5FST: the FAIR Speech Translation System for the IWSLT21 Multilingual Shared Task
In this paper, we describe our end-to-end multilingual speech translation system submitted to the IWSLT 2021 evaluation campaign on the Multilingual Speech Translation shared task. Our system is built by leveraging trans…
Transfer LearningTranslation