paper-with-me

Papers

Large-scale multilingual audio visual dubbing

2020-11-06 · Yi Yang, Brendan Shillingford, Yannis Assael, Miaosen Wang, Wendi Liu, Yutian Chen, Yu Zhang, Eren Sezener, Luis C. Cobo, Misha Denil, Yusuf Aytar, Nando de Freitas

We describe a system for large-scale audiovisual translation and dubbing, which translates videos from one language to another. The source language's speech content is transcribed to text, translated, and automatically synthesized into target language speech using the original speaker's voice. The visual content is translated by synthesizing lip movements for the speaker to match the translated audio, creating a seamless audiovisual experience in the target language. The audio and visual translation subsystems each contain a large-scale generic synthesis model trained on thousands of hours of data in the corresponding domain. These generic models are fine-tuned to a specific speaker before translation, either using an auxiliary corpus of data from the target speaker, or using the video to be translated itself as the input to the fine-tuning process. This report gives an architectural overview of the full system, as well as an in-depth discussion of the video dubbing component. The role of the audio and text components in relation to the full system is outlined, but their design is not discussed in detail. Translated and dubbed demo videos generated using our system can be viewed at https://www.youtube.com/playlist?list=PLSi232j2ZA6_1Exhof5vndzyfbxAhhEs5

📄 PDF Abstract BibTeX arXiv:2011.03530

Code (0)

등록된 구현이 없습니다.

Tasks

Translation

Similar Papers 제목 키워드 기반

JUST-DUB-IT: Video Dubbing via Joint Audio-Visual Diffusion

2026-01-29 · Anthony Chen, Naomi Ken Korem, Gal Zeevi, Tavi Halperin 외 arxiv

Audio-Visual Foundation Models, which are pretrained to jointly generate sound and visual content, have recently shown an unprecedented ability to model multi-modal generation and editing, opening new opportunities for d…

SyncVoice: Simple and Effective Automatic Video Dubbing with Vision-Augmented TTS

2025-11-23 · Kaidi Wang, Yi He, Wenhao Guan, Weijie Wu 외 arxiv

Automatic video dubbing aims to generate high-fidelity speech that is temporally aligned with visual content. However, existing methods still suffer from limited speech naturalness, insufficient audio-visual synchronizat…

Speech Synthesis

FunCineForge: A Unified Dataset Toolkit and Model for Zero-Shot Movie Dubbing in Diverse Cinematic Scenes

2026-01-21 · Jiaxuan Liu, Yang Xiang, Han Zhao, Xiangang Li 외 arxiv

Movie dubbing is the task of synthesizing speech from scripts conditioned on video scenes, requiring accurate lip sync, faithful timbre transfer, and proper modeling of character identity and emotion. However, existing m…

Instruction Following

Dubbing for Everyone: Data-Efficient Visual Dubbing using Neural Rendering Priors

2024-01-11 · Jack Saunders, Vinay Namboodiri

Visual dubbing is the process of generating lip motions of an actor in a video to synchronise with given audio. Recent advances have made progress towards this goal but have not been able to produce an approach suitable …

Neural Rendering

MINT: a Multi-modal Image and Narrative Text Dubbing Dataset for Foley Audio Content Planning and Generation

2024-06-15 · Ruibo Fu, Shuchen Shi, Hongming Guo, Tao Wang 외

Foley audio, critical for enhancing the immersive experience in multimedia content, faces significant challenges in the AI-generated content (AIGC) landscape. Despite advancements in AIGC technologies for text and image …

AudioCapsImage Generation