paper-with-me

홈 › Papers

Length Aware Speech Translation for Video Dubbing

2025-05-31 · Harveen Singh Chadha, Aswin Shanmugam Subramanian, Vikas Joshi, Shubham Bansal, Jian Xue, Rupeshkumar Mehta, Jinyu Li

In video dubbing, aligning translated audio with the source audio is a significant challenge. Our focus is on achieving this efficiently, tailored for real-time, on-device video dubbing scenarios. We developed a phoneme-based end-to-end length-sensitive speech translation (LSST) model, which generates translations of varying lengths short, normal, and long using predefined tags. Additionally, we introduced length-aware beam search (LABS), an efficient approach to generate translations of different lengths in a single decoding pass. This approach maintained comparable BLEU scores compared to a baseline without length awareness while significantly enhancing synchronization quality between source and target audio, achieving a mean opinion score (MOS) gain of 0.34 for Spanish and 0.65 for Korean, respectively.

📄 PDF Abstract BibTeX arXiv:2506.00740

Code (0)

등록된 구현이 없습니다.

Tasks

Translation

Methods 이 논문이 사용한 방법론

Focus 설명 없음

Similar Papers 제목 키워드 기반

VideoDubber: Machine Translation with Speech-Aware Length Control for Video Dubbing

2022-11-30 · Yihan Wu, Junliang Guo, Xu Tan, Chen Zhang 외

Video dubbing aims to translate the original speech in a film or television program into the speech in a target language, which can be achieved with a cascaded system consisting of speech recognition, machine translation…

Machine TranslationSentencespeech-recognitionSpeech Recognition+2

Dubbing in Practice: A Large Scale Study of Human Localization With Insights for Automatic Dubbing

2022-12-23 · William Brannon, Yogesh Virkar, Brian Thompson

We investigate how humans perform the task of dubbing video content from one language into another, leveraging a novel corpus of 319.57 hours of video from 54 professionally produced titles. This is the first such large-…

Translation

Isochrony-Aware Neural Machine Translation for Automatic Dubbing

2021-12-16 · Derek Tam, Surafel M. Lakew, Yogesh Virkar, Prashant Mathur 외

We introduce the task of isochrony-aware machine translation which aims at generating translations suitable for dubbing. Dubbing of a spoken sentence requires transferring the content as well as the speech-pause structur…

Machine TranslationSentenceTranslation

From Speech-to-Speech Translation to Automatic Dubbing

2020-01-19 · WS 2020 7 · Marcello Federico, Robert Enyedi, Roberto Barra-Chicote, Ritwik Giri 외

We present enhancements to a speech-to-speech translation pipeline in order to perform automatic dubbing. Our architecture features neural machine translation generating output of preferred length, prosodic alignment of …

Machine TranslationSpeech-to-Speech Translationtext-to-speechText to Speech+1

Jointly Optimizing Translations and Speech Timing to Improve Isochrony in Automatic Dubbing

2023-02-25 · Alexandra Chronopoulou, Brian Thompson, Prashant Mathur, Yogesh Virkar 외

Automatic dubbing (AD) is the task of translating the original speech in a video into target language speech. The new target language speech should satisfy isochrony; that is, the new speech should be time aligned with t…

Translation