paper-with-me

Papers

ConsistentAvatar: Learning to Diffuse Fully Consistent Talking Head Avatar with Temporal Guidance

2024-11-23 · Haijie Yang, Zhenyu Zhang, Hao Tang, Jianjun Qian, Jian Yang

Diffusion models have shown impressive potential on talking head generation. While plausible appearance and talking effect are achieved, these methods still suffer from temporal, 3D or expression inconsistency due to the error accumulation and inherent limitation of single-image generation ability. In this paper, we propose ConsistentAvatar, a novel framework for fully consistent and high-fidelity talking avatar generation. Instead of directly employing multi-modal conditions to the diffusion process, our method learns to first model the temporal representation for stability between adjacent frames. Specifically, we propose a Temporally-Sensitive Detail (TSD) map containing high-frequency feature and contours that vary significantly along the time axis. Using a temporal consistent diffusion module, we learn to align TSD of the initial result to that of the video frame ground truth. The final avatar is generated by a fully consistent diffusion module, conditioned on the aligned TSD, rough head normal, and emotion prompt embedding. We find that the aligned TSD, which represents the temporal patterns, constrains the diffusion process to generate temporally stable talking head. Further, its reliable guidance complements the inaccuracy of other conditions, suppressing the accumulated error while improving the consistency on various aspects. Extensive experiments demonstrate that ConsistentAvatar outperforms the state-of-the-art methods on the generated appearance, 3D, expression and temporal consistency. Project page: https://njust-yang.github.io/ConsistentAvatar.github.io/

📄 PDF Abstract BibTeX arXiv:2411.15436

Code (0)

등록된 구현이 없습니다.

Tasks

Image Generationsingle-image-generationTalking Head Generation

Methods 이 논문이 사용한 방법론

Diffusion Diffusion models generate samples by gradually removing noise from a signal, and their training objective can be expressed as a reweighted variational lower-bound…
ALIGN In the ALIGN method, visual and language representations are jointly trained from noisy image alt-text data. The image and text encoders are learned via contrastive loss…

Similar Papers 제목 키워드 기반

Warm Chat: Diffuse Emotion-aware Interactive Talking Head Avatar with Tree-Structured Guidance

2025-08-25 · Haijie Yang, Zhenyu Zhang, Hao Tang, Jianjun Qian 외 arxiv

Generative models have advanced rapidly, enabling impressive talking head generation that brings AI to life. However, most existing methods focus solely on one-way portrait animation. Even the few that support bidirectio…

Talking Head GenerationDialogue Generation

Diffused Heads: Diffusion Models Beat GANs on Talking-Face Generation

2023-01-06 · Michał Stypułkowski, Konstantinos Vougioukas, Sen He, Maciej Zięba 외

Talking face generation has historically struggled to produce head movements and natural facial expressions without guidance from additional reference videos. Recent developments in diffusion-based generative models allo…

Face GenerationTalking Face GenerationVideo Generation

Talk3D: High-Fidelity Talking Portrait Synthesis via Personalized 3D Generative Prior

2024-03-29 · Jaehoon Ko, Kyusun Cho, Joungbin Lee, Heeji Yoon 외

Recent methods for audio-driven talking head synthesis often optimize neural radiance fields (NeRF) on a monocular talking portrait video, leveraging its capability to render high-fidelity and 3D-consistent novel-view fr…

NeRF

EmoHead: Emotional Talking Head via Manipulating Semantic Expression Parameters

2025-03-25 · Xuli Shen, Hua Cai, Dingding Yu, Weilin Shen 외

Generating emotion-specific talking head videos from audio input is an important and complex challenge for human-machine interaction. However, emotion is highly abstract concept with ambiguous boundaries, and it necessit…

TAG

3D-Aware Talking-Head Video Motion Transfer

2023-11-05 · Haomiao Ni, Jiachen Liu, Yuan Xue, Sharon X. Huang

Motion transfer of talking-head videos involves generating a new video with the appearance of a subject video and the motion pattern of a driving video. Current methodologies primarily depend on a limited number of subje…

Novel View Synthesis