paper-with-me

Papers

DeepDubber-V1: Towards High Quality and Dialogue, Narration, Monologue Adaptive Movie Dubbing Via Multi-Modal Chain-of-Thoughts Reasoning Guidance

2025-03-31 · Junjie Zheng, Zihao Chen, Chaofan Ding, Xinhan Di

Current movie dubbing technology can generate the desired voice from a given speech prompt, ensuring good synchronization between speech and visuals while accurately conveying the intended emotions. However, in movie dubbing, key aspects such as adapting to different dubbing styles, handling dialogue, narration, and monologue effectively, and understanding subtle details like the age and gender of speakers, have not been well studied. To address this challenge, we propose a framework of multi-modal large language model. First, it utilizes multimodal Chain-of-Thought (CoT) reasoning methods on visual inputs to understand dubbing styles and fine-grained attributes. Second, it generates high-quality dubbing through large speech generation models, guided by multimodal conditions. Additionally, we have developed a movie dubbing dataset with CoT annotations. The evaluation results demonstrate a performance improvement over state-of-the-art methods across multiple datasets. In particular, for the evaluation metrics, the SPK-SIM and EMO-SIM increases from 82.48% to 89.74%, 66.24% to 78.88% for dubbing setting 2.0 on V2C Animation dataset, LSE-D and MCD-SL decreases from 14.79 to 14.63, 5.24 to 4.74 for dubbing setting 2.0 on Grid dataset, SPK-SIM increases from 64.03 to 83.42 and WER decreases from 52.69% to 23.20% for initial reasoning setting on proposed CoT-Movie-Dubbing dataset in the comparison with the state-of-the art models.

📄 PDF Abstract BibTeX arXiv:2503.23660

Code (0)

등록된 구현이 없습니다.

Tasks

Large Language Model

Similar Papers 제목 키워드 기반

FunCineForge: A Unified Dataset Toolkit and Model for Zero-Shot Movie Dubbing in Diverse Cinematic Scenes

2026-01-21 · Jiaxuan Liu, Yang Xiang, Han Zhao, Xiangang Li 외 arxiv

Movie dubbing is the task of synthesizing speech from scripts conditioned on video scenes, requiring accurate lip sync, faithful timbre transfer, and proper modeling of character identity and emotion. However, existing m…

Instruction Following

MonoTODia: Translating Monologue Requests to Task-Oriented Dialogues

2025-02-24 · Sebastian Steindl, Ulrich Schäfer, Bernd Ludwig

Data scarcity is one of the main problems when it comes to real-world applications of transformer-based models. This is especially evident for task-oriented dialogue (TOD) systems, which require specialized datasets, tha…

DialogueReason: Rule-Based RL Sparks Dialogue Reasoning in LLMs

2025-05-11 · Yubo Shu, Zhewei Huang, Xin Wu, Chen Hu 외

We propose DialogueReason, a reasoning paradigm that uncovers the lost roles in monologue-style reasoning models, aiming to boost diversity and coherency of the reasoning process. Recent advances in RL-based large reason…

DiversityMath

MM-MovieDubber: Towards Multi-Modal Learning for Multi-Modal Movie Dubbing

2025-05-22 · Junjie Zheng, Zihao Chen, Chaofan Ding, Yunming Liang 외

Current movie dubbing technology can produce the desired speech using a reference voice and input video, maintaining perfect synchronization with the visuals while effectively conveying the intended emotions. However, cr…

Language ModelingLanguage Modelling

Dialogue Coherence Assessment Without Explicit Dialogue Act Labels

2019-08-22 · ACL 2020 6 · Mohsen Mesgar, Sebastian Bücker, Iryna Gurevych

Recent dialogue coherence models use the coherence features designed for monologue texts, e.g. nominal entities, to represent utterances and then explicitly augment them with dialogue-relevant features, e.g., dialogue ac…

Multi-Task Learning