paper-with-me

Papers

InfiniteTalk: Audio-driven Video Generation for Sparse-Frame Video Dubbing

2025-08-19 · Shaoshu Yang, Zhe Kong, Feng Gao, Meng Cheng, Xiangyu Liu, Yong Zhang, Zhuoliang Kang, Wenhan Luo, Xunliang Cai, Ran He, Xiaoming Wei arxiv

Recent breakthroughs in video AIGC have ushered in a transformative era for audio-driven human animation. However, conventional video dubbing techniques remain constrained to mouth region editing, resulting in discordant facial expressions and body gestures that compromise viewer immersion. To overcome this limitation, we introduce sparse-frame video dubbing, a novel paradigm that strategically preserves reference keyframes to maintain identity, iconic gestures, and camera trajectories while enabling holistic, audio-synchronized full-body motion editing. Through critical analysis, we identify why naive image-to-video models fail in this task, particularly their inability to achieve adaptive conditioning. Addressing this, we propose InfiniteTalk, a streaming audio-driven generator designed for infinite-length long sequence dubbing. This architecture leverages temporal context frames for seamless inter-chunk transitions and incorporates a simple yet effective sampling strategy that optimizes control strength via fine-grained reference frame positioning. Comprehensive evaluations on HDTF, CelebV-HQ, and EMTD datasets demonstrate state-of-the-art performance. Quantitative metrics confirm superior visual realism, emotional coherence, and full-body motion synchronization.

📄 PDF Abstract BibTeX arXiv:2508.14033

Code (0)

등록된 구현이 없습니다.

Tasks

Video Generation

Similar Papers 제목 키워드 기반

TalkingMachines: Real-Time Audio-Driven FaceTime-Style Video via Autoregressive Diffusion Models

2025-06-03 · Chetwin Low, Weimin WANG

In this paper, we present TalkingMachines -- an efficient framework that transforms pretrained video generation models into real-time, audio-driven character animators. TalkingMachines enables natural conversational expe…

DecoderKnowledge DistillationLanguage ModelingLanguage Modelling+2

AudCast: Audio-Driven Human Video Generation by Cascaded Diffusion Transformers

2025-03-25 · CVPR 2025 1 · Jiazhi Guan, Kaisiyuan Wang, Zhiliang Xu, Quanwei Yang 외

Despite the recent progress of audio-driven video generation, existing methods mostly focus on driving facial movements, leading to non-coherent head and body dynamics. Moving forward, it is desirable yet challenging to …

Video Generation

Input-Aware Sparse Attention for Real-Time Co-Speech Video Generation

2025-10-02 · Beijia Lu, Ziyi Chen, Jing Xiao, Jun-Yan Zhu arxiv

Diffusion models can synthesize realistic co-speech video from audio for various applications, such as video creation and virtual agents. However, existing diffusion-based methods are slow due to numerous denoising steps…

Video Generation

OmniAvatar: Efficient Audio-Driven Avatar Video Generation with Adaptive Body Animation

2025-06-23 · Qijun Gan, Ruizi Yang, Jianke Zhu, Shaofei Xue 외

Significant progress has been made in audio-driven human animation, while most existing methods focus mainly on facial movements, limiting their ability to create full-body animations with natural synchronization and flu…

Human AnimationVideo Generation

TalkVerse: Democratizing Minute-Long Audio-Driven Video Generation

2025-12-16 · Zhenzhi Wang, Jian Wang, Ke Ma, Dahua Lin 외 arxiv

We introduce TalkVerse, a large-scale, open corpus for single-person, audio-driven talking video generation designed to enable fair, reproducible comparison across methods. While current state-of-the-art systems rely on …

Video Generation