paper-with-me

Papers

DreamHead: Learning Spatial-Temporal Correspondence via Hierarchical Diffusion for Audio-driven Talking Head Synthesis

2024-09-16 · Fa-Ting Hong, Yunfei Liu, Yu Li, Changyin Zhou, Fei Yu, Dan Xu

Audio-driven talking head synthesis strives to generate lifelike video portraits from provided audio. The diffusion model, recognized for its superior quality and robust generalization, has been explored for this task. However, establishing a robust correspondence between temporal audio cues and corresponding spatial facial expressions with diffusion models remains a significant challenge in talking head generation. To bridge this gap, we present DreamHead, a hierarchical diffusion framework that learns spatial-temporal correspondences in talking head synthesis without compromising the model's intrinsic quality and adaptability.~DreamHead learns to predict dense facial landmarks from audios as intermediate signals to model the spatial and temporal correspondences.~Specifically, a first hierarchy of audio-to-landmark diffusion is first designed to predict temporally smooth and accurate landmark sequences given audio sequence signals. Then, a second hierarchy of landmark-to-image diffusion is further proposed to produce spatially consistent facial portrait videos, by modeling spatial correspondences between the dense facial landmark and appearance. Extensive experiments show that proposed DreamHead can effectively learn spatial-temporal consistency with the designed hierarchical diffusion and produce high-fidelity audio-driven talking head videos for multiple identities.

📄 PDF Abstract BibTeX arXiv:2409.10281

Code (0)

등록된 구현이 없습니다.

Tasks

Talking Head Generation

Methods 이 논문이 사용한 방법론

Diffusion Diffusion models generate samples by gradually removing noise from a signal, and their training objective can be expressed as a reweighted variational lower-bound…

Similar Papers 제목 키워드 기반

Disentangled Diffusion-Based 3D Human Pose Estimation with Hierarchical Spatial and Temporal Denoiser

2024-03-07 · Qingyuan Cai, Xuecai Hu, Saihui Hou, Li Yao 외

Recently, diffusion-based methods for monocular 3D human pose estimation have achieved state-of-the-art (SOTA) performance by directly regressing the 3D joint coordinates from the 2D pose sequence. Although some methods …

3D Human Pose EstimationDisentanglementMonocular 3D Human Pose EstimationMulti-Hypotheses 3D Human Pose Estimation+1

FRESCO: Spatial-Temporal Correspondence for Zero-Shot Video Translation

2024-03-19 · CVPR 2024 1 · Shuai Yang, Yifan Zhou, Ziwei Liu, Chen Change Loy

The remarkable efficacy of text-to-image diffusion models has motivated extensive exploration of their potential application in video domains. Zero-shot methods seek to extend image diffusion models to videos without nec…

Translationvalid

Zero-Shot Video Translation and Editing with Frame Spatial-Temporal Correspondence

2025-12-03 · Shuai Yang, Junxin Lin, Yifan Zhou, Ziwei Liu 외 arxiv

The remarkable success in text-to-image diffusion models has motivated extensive investigation of their potential for video applications. Zero-shot techniques aim to adapt image diffusion models for videos without requir…

GMS-CAVP: Improving Audio-Video Correspondence with Multi-Scale Contrastive and Generative Pretraining

2026-01-27 · Shentong Mo, Zehua Chen, Jun Zhu arxiv

Recent advances in video-audio (V-A) understanding and generation have increasingly relied on joint V-A embeddings, which serve as the foundation for tasks such as cross-modal retrieval and generation. While prior method…

Cross-Modal RetrievalContrastive Learning

Unsupervised Foggy Scene Understanding via Self Spatial-Temporal Label Diffusion

2022-06-10 · Liang Liao, WenYi Chen, Jing Xiao, Zheng Wang 외

Understanding foggy image sequence in the driving scenes is critical for autonomous driving, but it remains a challenging task due to the difficulty in collecting and annotating real-world images of adverse weather. Rece…

Autonomous DrivingDomain AdaptationPseudo LabelScene Understanding+3