paper-with-me

홈 › Papers

Talking Head Generation with Probabilistic Audio-to-Visual Diffusion Priors

2022-12-07 · ICCV 2023 1 · Zhentao Yu, Zixin Yin, Deyu Zhou, Duomin Wang, Finn Wong, Baoyuan Wang

In this paper, we introduce a simple and novel framework for one-shot audio-driven talking head generation. Unlike prior works that require additional driving sources for controlled synthesis in a deterministic manner, we instead probabilistically sample all the holistic lip-irrelevant facial motions (i.e. pose, expression, blink, gaze, etc.) to semantically match the input audio while still maintaining both the photo-realism of audio-lip synchronization and the overall naturalness. This is achieved by our newly proposed audio-to-visual diffusion prior trained on top of the mapping between audio and disentangled non-lip facial representations. Thanks to the probabilistic nature of the diffusion prior, one big advantage of our framework is it can synthesize diverse facial motion sequences given the same audio clip, which is quite user-friendly for many real applications. Through comprehensive evaluations on public benchmarks, we conclude that (1) our diffusion prior outperforms auto-regressive prior significantly on almost all the concerned metrics; (2) our overall system is competitive with prior works in terms of audio-lip synchronization but can effectively sample rich and natural-looking lip-irrelevant facial motions while still semantically harmonized with the audio input.

📄 PDF Abstract BibTeX arXiv:2212.04248

Code (0)

등록된 구현이 없습니다.

Tasks

Talking Head Generation

Methods 이 논문이 사용한 방법론

Diffusion Diffusion models generate samples by gradually removing noise from a signal, and their training objective can be expressed as a reweighted variational lower-bound…

Similar Papers 제목 키워드 기반

FONT: Flow-guided One-shot Talking Head Generation with Natural Head Motions

2023-03-31 · Jin Liu, Xi Wang, Xiaomeng Fu, Yesheng Chai 외

One-shot talking head generation has received growing attention in recent years, with various creative and practical applications. An ideal natural and vivid generated talking head video should contain natural head pose …

DiversityPose PredictionTalking Head GenerationUnsupervised Keypoints

OmniTalker: Real-Time Text-Driven Talking Head Generation with In-Context Audio-Visual Style Replication

2025-04-03 · Zhongjian Wang, Peng Zhang, Jinwei Qi, Guangyuan Wang Sheng Xu 외

Recent years have witnessed remarkable advances in talking head generation, owing to its potential to revolutionize the human-AI interaction from text interfaces into realistic video chats. However, research on text-driv…

Talking Head GenerationVideo Synchronization

Dual Audio-Centric Modality Coupling for Talking Head Generation

2025-03-26 · Ao Fu, Ziqi Ni, Yi Zhou

The generation of audio-driven talking head videos is a key challenge in computer vision and graphics, with applications in virtual avatars and digital media. Traditional approaches often struggle with capturing the comp…

NeRFTalking Head Generationtext-to-speechText to Speech

Expressive Talking Head Generation With Granular Audio-Visual Control

2022-01-01 · CVPR 2022 1 · Borong Liang, Yan Pan, Zhizhi Guo, Hang Zhou 외

Generating expressive talking heads is essential for creating virtual humans. However, existing one- or few-shot methods focus on lip-sync and head motion, ignoring the emotional expressions that make talking faces r…

Talking Head Generation

VividTalk: One-Shot Audio-Driven Talking Head Generation Based on 3D Hybrid Prior

2023-12-04 · Xusen Sun, Longhao Zhang, Hao Zhu, Peng Zhang 외

Audio-driven talking head generation has drawn much attention in recent years, and many efforts have been made in lip-sync, expressive facial expressions, natural head pose generation, and high video quality. However, no…

Talking Head Generation