paper-with-me

홈 › Papers

Fine-grained Video Dubbing Duration Alignment with Segment Supervised Preference Optimization

2025-08-12 · Chaoqun Cui, Liangbin Huang, Shijing Wang, Zhe Tong, Zhaolong Huang, Xiao Zeng, Xiaofeng Liu arxiv

Video dubbing aims to translate original speech in visual media programs from the source language to the target language, relying on neural machine translation and text-to-speech technologies. Due to varying information densities across languages, target speech often mismatches the source speech duration, causing audio-video synchronization issues that significantly impact viewer experience. In this study, we approach duration alignment in LLM-based video dubbing machine translation as a preference optimization problem. We propose the Segment Supervised Preference Optimization (SSPO) method, which employs a segment-wise sampling strategy and fine-grained loss to mitigate duration mismatches between source and target lines. Experimental results demonstrate that SSPO achieves superior performance in duration alignment tasks.

📄 PDF Abstract BibTeX arXiv:2508.08550

Code (0)

등록된 구현이 없습니다.

Tasks

Machine Translation

Similar Papers 제목 키워드 기반

CoSyncDiT: Cognitive Synchronous Diffusion Transformer for Movie Dubbing

2026-04-14 · Gaoxiang Cong, Liang Li, Jiaxin Ye, Zhedong Zhang 외 arxiv

Movie dubbing aims to synthesize speech that preserves the vocal identity of a reference audio while synchronizing with the lip movements in a target video. Existing methods fail to achieve precise lip-sync and lack natu…

DubWise: Video-Guided Speech Duration Control in Multimodal LLM-based Text-to-Speech for Dubbing

2024-06-13 · Neha Sahipjohn, Ashishkumar Gudmalwar, Nirmesh Shah, Pankaj Wasnik 외

Audio-visual alignment after dubbing is a challenging research problem. To this end, we propose a novel method, DubWise Multi-modal Large Language Model (LLM)-based Text-to-Speech (TTS), which can control the speech dura…

Language ModelingLanguage ModellingLarge Language Modeltext-to-speech+2

Identity-Preserving Video Dubbing Using Motion Warping

2025-01-08 · Runzhen Liu, Qinjie Lin, Yunfei Liu, Lijian Lin 외

Video dubbing aims to synthesize realistic, lip-synced videos from a reference video and a driving audio signal. Although existing methods can accurately generate mouth shapes driven by audio, they often fail to preserve…

VideoDubber: Machine Translation with Speech-Aware Length Control for Video Dubbing

2022-11-30 · Yihan Wu, Junliang Guo, Xu Tan, Chen Zhang 외

Video dubbing aims to translate the original speech in a film or television program into the speech in a target language, which can be achieved with a cascaded system consisting of speech recognition, machine translation…

Machine TranslationSentencespeech-recognitionSpeech Recognition+2

DiFlowDubber: Discrete Flow Matching for Automated Video Dubbing via Cross-Modal Alignment and Synchronization

2026-03-15 · Ngoc-Son Nguyen, Thanh V. T. Tran, Jeongsoo Choi, Hieu-Nghia Huynh-Nguyen 외 arxiv

Video dubbing requires content accuracy, expressive prosody, high-quality acoustics, and precise lip synchronization, yet existing approaches struggle on all four fronts. To address these issues, we propose DiFlowDubber,…