paper-with-me

Papers

Dubbing in Practice: A Large Scale Study of Human Localization With Insights for Automatic Dubbing

2022-12-23 · William Brannon, Yogesh Virkar, Brian Thompson

We investigate how humans perform the task of dubbing video content from one language into another, leveraging a novel corpus of 319.57 hours of video from 54 professionally produced titles. This is the first such large-scale study we are aware of. The results challenge a number of assumptions commonly made in both qualitative literature on human dubbing and machine-learning literature on automatic dubbing, arguing for the importance of vocal naturalness and translation quality over commonly emphasized isometric (character length) and lip-sync constraints, and for a more qualified view of the importance of isochronic (timing) constraints. We also find substantial influence of the source-side audio on human dubs through channels other than the words of the translation, pointing to the need for research on ways to preserve speech characteristics, as well as semantic transfer such as emphasis/emotion, in automatic dubbing systems.

📄 PDF Abstract BibTeX arXiv:2212.12137

Code (1)

amazon-science/iwslt-autodub-task

Tasks

Translation

Methods 이 논문이 사용한 방법론

AWARE We propose to theoretically and empirically examine the effect of incorporating weighting schemes into walk-aggregating GNNs. To this end, we propose a simple, interpretable, and…

Similar Papers 제목 키워드 기반

FunCineForge: A Unified Dataset Toolkit and Model for Zero-Shot Movie Dubbing in Diverse Cinematic Scenes

2026-01-21 · Jiaxuan Liu, Yang Xiang, Han Zhao, Xiangang Li 외 arxiv

Movie dubbing is the task of synthesizing speech from scripts conditioned on video scenes, requiring accurate lip sync, faithful timbre transfer, and proper modeling of character identity and emotion. However, existing m…

Instruction Following

SyncVoice: Simple and Effective Automatic Video Dubbing with Vision-Augmented TTS

2025-11-23 · Kaidi Wang, Yi He, Wenhao Guan, Weijie Wu 외 arxiv

Automatic video dubbing aims to generate high-fidelity speech that is temporally aligned with visual content. However, existing methods still suffer from limited speech naturalness, insufficient audio-visual synchronizat…

Speech Synthesis

Technology Pipeline for Large Scale Cross-Lingual Dubbing of Lecture Videos into Multiple Indian Languages

2022-11-01 · Anusha Prakash, Arun Kumar, Ashish Seth, Bhagyashree Mukherjee 외

Cross-lingual dubbing of lecture videos requires the transcription of the original audio, correction and removal of disfluencies, domain term discovery, text-to-text translation into the target language, chunking of text…

ChunkingRhythmSpeech Synthesistext-to-speech+2

Expressive Machine Dubbing Through Phrase-level Cross-lingual Prosody Transfer

2023-06-20 · Jakub Swiatkowski, Duo Wang, Mikolaj Babianski, Giuseppe Coccia 외

Speech generation for machine dubbing adds complexity to conventional Text-To-Speech solutions as the generated output is required to match the expressiveness, emotion and speaking rate of the source content. Capturing a…

text-to-speechText to Speech

MM-MovieDubber: Towards Multi-Modal Learning for Multi-Modal Movie Dubbing

2025-05-22 · Junjie Zheng, Zihao Chen, Chaofan Ding, Yunming Liang 외

Current movie dubbing technology can produce the desired speech using a reference voice and input video, maintaining perfect synchronization with the visuals while effectively conveying the intended emotions. However, cr…

Language ModelingLanguage Modelling