paper-with-me

Papers

ATL-Diff: Audio-Driven Talking Head Generation with Early Landmarks-Guide Noise Diffusion

2025-07-17 · Hoang-Son Vo, Quang-Vinh Nguyen, Seungwon Kim, Hyung-Jeong Yang, Soonja Yeom, Soo-Hyung Kim arxiv

Audio-driven talking head generation requires precise synchronization between facial animations and audio signals. This paper introduces ATL-Diff, a novel approach addressing synchronization limitations while reducing noise and computational costs. Our framework features three key components: a Landmark Generation Module converting audio to facial landmarks, a Landmarks-Guide Noise approach that decouples audio by distributing noise according to landmarks, and a 3D Identity Diffusion network preserving identity characteristics. Experiments on MEAD and CREMA-D datasets demonstrate that ATL-Diff outperforms state-of-the-art methods across all metrics. Our approach achieves near real-time processing with high-quality animations, computational efficiency, and exceptional preservation of facial nuances. This advancement offers promising applications for virtual assistants, education, medical communication, and digital platforms. The source code is available at: \href{https://github.com/sonvth/ATL-Diff}{https://github.com/sonvth/ATL-Diff}

📄 PDF Abstract BibTeX arXiv:2507.12804

Code (0)

등록된 구현이 없습니다.

Tasks

Computational EfficiencyTalking Head Generation

Similar Papers 제목 키워드 기반

AI killed the video star. Audio-driven diffusion model for expressive talking head generation

2025-11-27 · Baptiste Chopin, Tashvik Dhamija, Pranav Balaji, Yaohui Wang 외 arxiv

We propose Dimitra++, a novel framework for audio-driven talking head generation, streamlined to learn lip motion, facial expression, as well as head pose motion. Specifically, we propose a conditional Motion Diffusion T…

Talking Head Generation

DreamHead: Learning Spatial-Temporal Correspondence via Hierarchical Diffusion for Audio-driven Talking Head Synthesis

2024-09-16 · Fa-Ting Hong, Yunfei Liu, Yu Li, Changyin Zhou 외

Audio-driven talking head synthesis strives to generate lifelike video portraits from provided audio. The diffusion model, recognized for its superior quality and robust generalization, has been explored for this task. H…

Talking Head Generation

READ: Real-time and Efficient Asynchronous Diffusion for Audio-driven Talking Head Generation

2025-08-05 · Haotian Wang, Yuzhe Weng, Jun Du, Haoran Xu 외 arxiv

The introduction of diffusion models has brought significant advances to the field of audio-driven talking head generation. However, the extremely slow inference speed severely limits the practical implementation of diff…

Talking Head Generation

StyleTalker: One-shot Style-based Audio-driven Talking Head Video Generation

2022-08-23 · Dongchan Min, Minyoung Song, Eunji Ko, Sung Ju Hwang

We propose StyleTalker, a novel audio-driven talking head generation model that can synthesize a video of a talking person from a single reference image with accurately audio-synced lip shapes, realistic head poses, and …

Talking Head GenerationVideo Generation

Dimitra: Audio-driven Diffusion model for Expressive Talking Head Generation

2025-02-24 · Baptiste Chopin, Tashvik Dhamija, Pranav Balaji, Yaohui Wang 외

We propose Dimitra, a novel framework for audio-driven talking head generation, streamlined to learn lip motion, facial expression, as well as head pose motion. Specifically, we train a conditional Motion Diffusion Trans…

Talking Head Generation