paper-with-me

홈 › Papers

Dub-S2ST: Textless Speech-to-Speech Translation for Seamless Dubbing

2025-05-27 · Jeongsoo Choi, Jaehun Kim, Joon Son Chung

This paper introduces a cross-lingual dubbing system that translates speech from one language to another while preserving key characteristics such as duration, speaker identity, and speaking speed. Despite the strong translation quality of existing speech translation approaches, they often overlook the transfer of speech patterns, leading to mismatches with source speech and limiting their suitability for dubbing applications. To address this, we propose a discrete diffusion-based speech-to-unit translation model with explicit duration control, enabling time-aligned translation. We then synthesize speech based on the predicted units and source identity with a conditional flow matching model. Additionally, we introduce a unit-based speed adaptation mechanism that guides the translation model to produce speech at a rate consistent with the source, without relying on any text. Extensive experiments demonstrate that our framework generates natural and fluent translations that align with the original speech's duration and speaking pace, while achieving competitive translation performance.

📄 PDF Abstract BibTeX arXiv:2505.20899

Code (0)

등록된 구현이 없습니다.

Tasks

Speech-to-Speech TranslationTranslation

Methods 이 논문이 사용한 방법론

SPEED The monocular depth estimation (MDE) is the task of estimating depth from a single frame. This information is an essential knowledge in many computer vision tasks such as scene…
ALIGN In the ALIGN method, visual and language representations are jointly trained from noisy image alt-text data. The image and text encoders are learned via contrastive loss…

Similar Papers 제목 키워드 기반

Machine Translation Verbosity Control for Automatic Dubbing

2021-10-08 · Surafel M. Lakew, Marcello Federico, Yue Wang, Cuong Hoang 외

Automatic dubbing aims at seamlessly replacing the speech in a video document with synthetic speech in a different language. The task implies many challenges, one of which is generating translations that not only convey …

Machine TranslationTranslation

From Speech-to-Speech Translation to Automatic Dubbing

2020-01-19 · WS 2020 7 · Marcello Federico, Robert Enyedi, Roberto Barra-Chicote, Ritwik Giri 외

We present enhancements to a speech-to-speech translation pipeline in order to perform automatic dubbing. Our architecture features neural machine translation generating output of preferred length, prosodic alignment of …

Machine TranslationSpeech-to-Speech Translationtext-to-speechText to Speech+1

Textless Speech-to-Speech Translation on Real Data

2021-12-15 · NAACL 2022 7 · Ann Lee, Hongyu Gong, Paul-Ambroise Duquenne, Holger Schwenk 외

We present a textless speech-to-speech translation (S2ST) system that can translate speech from one language into another language and can be built without the need of any text data. Different from existing work in the l…

Speech-to-Speech TranslationTranslation

Textless Speech-to-Speech Translation With Limited Parallel Data

2023-05-24 · Anuj Diwan, Anirudh Srinivasan, David Harwath, Eunsol Choi

Existing speech-to-speech translation (S2ST) models fall into two camps: they either leverage text as an intermediate step or require hundreds of hours of parallel speech data. Both approaches are incompatible with textl…

Automatic Speech RecognitionDenoisingLanguage ModellingMachine Translation+4

VideoDubber: Machine Translation with Speech-Aware Length Control for Video Dubbing

2022-11-30 · Yihan Wu, Junliang Guo, Xu Tan, Chen Zhang 외

Video dubbing aims to translate the original speech in a film or television program into the speech in a target language, which can be achieved with a cascaded system consisting of speech recognition, machine translation…

Machine TranslationSentencespeech-recognitionSpeech Recognition+2