paper-with-me

홈 › Papers

DUB: Discrete Unit Back-translation for Speech Translation

2023-05-19 · Dong Zhang, Rong Ye, Tom Ko, Mingxuan Wang, Yaqian Zhou

How can speech-to-text translation (ST) perform as well as machine translation (MT)? The key point is to bridge the modality gap between speech and text so that useful MT techniques can be applied to ST. Recently, the approach of representing speech with unsupervised discrete units yields a new way to ease the modality problem. This motivates us to propose Discrete Unit Back-translation (DUB) to answer two questions: (1) Is it better to represent speech with discrete units than with continuous features in direct ST? (2) How much benefit can useful MT techniques bring to ST? With DUB, the back-translation technique can successfully be applied on direct ST and obtains an average boost of 5.5 BLEU on MuST-C En-De/Fr/Es. In the low-resource language scenario, our method achieves comparable performance to existing methods that rely on large-scale external data. Code and models are available at https://github.com/0nutation/DUB.

📄 PDF Abstract BibTeX arXiv:2305.11411

Code (1)

0nutation/dub 공식 구현 pytorch

Tasks

Machine TranslationSpeech-to-TextSpeech-to-Text TranslationTranslation

Similar Papers 제목 키워드 기반

Direct Punjabi to English speech translation using discrete units

2024-02-25 · Prabhjot Kaur, L. Andrew M. Bush, Weisong Shi

Speech-to-speech translation is yet to reach the same level of coverage as text-to-text translation systems. The current speech technology is highly limited in its coverage of over 7000 languages spoken worldwide, leavin…

Speech-to-Speech TranslationSpeech-to-TextTranslation

Back Translation for Speech-to-text Translation Without Transcripts

2023-05-15 · Qingkai Fang, Yang Feng

The success of end-to-end speech-to-text translation (ST) is often achieved by utilizing source transcripts, e.g., by pre-training with automatic speech recognition (ASR) and machine translation (MT) tasks, or by introdu…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)de-enMachine Translation+5

TransFace: Unit-Based Audio-Visual Speech Synthesizer for Talking Head Translation

2023-12-23 · Xize Cheng, Rongjie Huang, Linjun Li, Tao Jin 외

Direct speech-to-speech translation achieves high-quality results through the introduction of discrete units obtained from self-supervised learning. This approach circumvents delays and cascading errors associated with m…

es-enfr-enSelf-Supervised LearningSpeech-to-Speech Translation+1

DiffS2UT: A Semantic Preserving Diffusion Model for Textless Direct Speech-to-Speech Translation

2023-10-26 · Yongxin Zhu, Zhujin Gao, Xinyuan Zhou, Zhongyi Ye 외

While Diffusion Generative Models have achieved great success on image generation tasks, how to efficiently and effectively incorporate them into speech generation especially translation tasks remains a non-trivial probl…

Image GenerationSpeech-to-Speech TranslationTranslation

Direct Simultaneous Speech-to-Speech Translation with Variational Monotonic Multihead Attention

2021-10-15 · Xutai Ma, Hongyu Gong, Danni Liu, Ann Lee 외

We present a direct simultaneous speech-to-speech translation (Simul-S2ST) model, Furthermore, the generation of translation is independent from intermediate text representations. Our approach leverages recent progress o…

Simultaneous Speech-to-Speech TranslationSpeech SynthesisSpeech-to-Speech TranslationTranslation