paper-with-me

Papers

JUST-DUB-IT: Video Dubbing via Joint Audio-Visual Diffusion

2026-01-29 · Anthony Chen, Naomi Ken Korem, Gal Zeevi, Tavi Halperin, Matan Ben Yosef, Urska Jelercic, Ofir Bibi, Or Patashnik, Daniel Cohen-Or arxiv

Audio-Visual Foundation Models, which are pretrained to jointly generate sound and visual content, have recently shown an unprecedented ability to model multi-modal generation and editing, opening new opportunities for downstream tasks. Among these tasks, video dubbing could greatly benefit from such priors, yet most existing solutions still rely on complex, task-specific pipelines that struggle in real-world settings. In this work, we introduce a single-model approach that adapts a foundational audio-video diffusion model for video-to-video dubbing via a lightweight LoRA. The LoRA enables the model to condition on an input audio-video while jointly generating translated audio and synchronized facial motion. To train this LoRA, we leverage the generative model itself to synthesize paired multilingual videos of the same speaker. Specifically, we generate multilingual videos with language switches within a single clip, and then inpaint the face and audio in each half to match the language of the other half. By leveraging the rich generative prior of the audio-visual model, our approach preserves speaker identity and lip synchronization while remaining robust to complex motion and real-world dynamics. We demonstrate that our approach produces high-quality dubbed videos with improved visual fidelity, lip synchronization, and robustness compared to existing dubbing pipelines.

📄 PDF Abstract BibTeX arXiv:2601.22143

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Dubbing for Everyone: Data-Efficient Visual Dubbing using Neural Rendering Priors

2024-01-11 · Jack Saunders, Vinay Namboodiri

Visual dubbing is the process of generating lip motions of an actor in a video to synchronise with given audio. Recent advances have made progress towards this goal but have not been able to produce an approach suitable …

Neural Rendering

SyncVoice: Simple and Effective Automatic Video Dubbing with Vision-Augmented TTS

2025-11-23 · Kaidi Wang, Yi He, Wenhao Guan, Weijie Wu 외 arxiv

Automatic video dubbing aims to generate high-fidelity speech that is temporally aligned with visual content. However, existing methods still suffer from limited speech naturalness, insufficient audio-visual synchronizat…

Speech Synthesis

Large-scale multilingual audio visual dubbing

2020-11-06 · Yi Yang, Brendan Shillingford, Yannis Assael, Miaosen Wang 외

We describe a system for large-scale audiovisual translation and dubbing, which translates videos from one language to another. The source language's speech content is transcribed to text, translated, and automatically s…

Translation

Identity-Preserving Video Dubbing Using Motion Warping

2025-01-08 · Runzhen Liu, Qinjie Lin, Yunfei Liu, Lijian Lin 외

Video dubbing aims to synthesize realistic, lip-synced videos from a reference video and a driving audio signal. Although existing methods can accurately generate mouth shapes driven by audio, they often fail to preserve…

InfiniteTalk: Audio-driven Video Generation for Sparse-Frame Video Dubbing

2025-08-19 · Shaoshu Yang, Zhe Kong, Feng Gao, Meng Cheng 외 arxiv

Recent breakthroughs in video AIGC have ushered in a transformative era for audio-driven human animation. However, conventional video dubbing techniques remain constrained to mouth region editing, resulting in discordant…

Video Generation