paper-with-me

Papers

AudCast: Audio-Driven Human Video Generation by Cascaded Diffusion Transformers

2025-03-25 · CVPR 2025 1 · Jiazhi Guan, Kaisiyuan Wang, Zhiliang Xu, Quanwei Yang, Yasheng Sun, Shengyi He, Borong Liang, Yukang Cao, YingYing Li, Haocheng Feng, Errui Ding, Jingdong Wang, Youjian Zhao, Hang Zhou, Ziwei Liu

Despite the recent progress of audio-driven video generation, existing methods mostly focus on driving facial movements, leading to non-coherent head and body dynamics. Moving forward, it is desirable yet challenging to generate holistic human videos with both accurate lip-sync and delicate co-speech gestures w.r.t. given audio. In this work, we propose AudCast, a generalized audio-driven human video generation framework adopting a cascade Diffusion-Transformers (DiTs) paradigm, which synthesizes holistic human videos based on a reference image and a given audio. 1) Firstly, an audio-conditioned Holistic Human DiT architecture is proposed to directly drive the movements of any human body with vivid gesture dynamics. 2) Then to enhance hand and face details that are well-knownly difficult to handle, a Regional Refinement DiT leverages regional 3D fitting as the bridge to reform the signals, producing the final results. Extensive experiments demonstrate that our framework generates high-fidelity audio-driven holistic human videos with temporal coherence and fine facial and hand details. Resources can be found at https://guanjz20.github.io/projects/AudCast.

📄 PDF Abstract BibTeX arXiv:2503.19824

Code (0)

등록된 구현이 없습니다.

Tasks

Video Generation

Methods 이 논문이 사용한 방법론

Focus 설명 없음

Similar Papers 제목 키워드 기반

OmniAvatar: Efficient Audio-Driven Avatar Video Generation with Adaptive Body Animation

2025-06-23 · Qijun Gan, Ruizi Yang, Jianke Zhu, Shaofei Xue 외

Significant progress has been made in audio-driven human animation, while most existing methods focus mainly on facial movements, limiting their ability to create full-body animations with natural synchronization and flu…

Human AnimationVideo Generation

OmniHuman-1: Rethinking the Scaling-Up of One-Stage Conditioned Human Animation Models

2025-02-03 · Gaojie Lin, Jianwen Jiang, Jiaqi Yang, Zerong Zheng 외

End-to-end human animation, such as audio-driven talking human generation, has undergone notable advancements in the recent few years. However, existing methods still struggle to scale up as large general video generatio…

Human AnimationHuman-Object Interaction DetectionMotion GenerationVideo Generation

Audio-Driven Co-Speech Gesture Video Generation

2022-12-05 · Xian Liu, Qianyi Wu, Hang Zhou, Yuanqi Du 외

Co-speech gesture is crucial for human-machine interaction and digital entertainment. While previous works mostly map speech audio to human skeletons (e.g., 2D keypoints), directly generating speakers' gestures in the im…

Video Generation

SpA2V: Harnessing Spatial Auditory Cues for Audio-driven Spatially-aware Video Generation

2025-08-01 · Kien T. Pham, Yingqing He, Yazhou Xing, Qifeng Chen 외 arxiv

Audio-driven video generation aims to synthesize realistic videos that align with input audio recordings, akin to the human ability to visualize scenes from auditory input. However, existing approaches predominantly focu…

Video Generation

TalkVerse: Democratizing Minute-Long Audio-Driven Video Generation

2025-12-16 · Zhenzhi Wang, Jian Wang, Ke Ma, Dahua Lin 외 arxiv

We introduce TalkVerse, a large-scale, open corpus for single-person, audio-driven talking video generation designed to enable fair, reproducible comparison across methods. While current state-of-the-art systems rely on …

Video Generation