paper-with-me

홈 › Papers

Supervising 3D Talking Head Avatars with Analysis-by-Audio-Synthesis

2025-04-18 · Radek Daněček, Carolin Schmitt, Senya Polikovsky, Michael J. Black

In order to be widely applicable, speech-driven 3D head avatars must articulate their lips in accordance with speech, while also conveying the appropriate emotions with dynamically changing facial expressions. The key problem is that deterministic models produce high-quality lip-sync but without rich expressions, whereas stochastic models generate diverse expressions but with lower lip-sync quality. To get the best of both, we seek a stochastic model with accurate lip-sync. To that end, we develop a new approach based on the following observation: if a method generates realistic 3D lip motions, it should be possible to infer the spoken audio from the lip motion. The inferred speech should match the original input audio, and erroneous predictions create a novel supervision signal for training 3D talking head avatars with accurate lip-sync. To demonstrate this effect, we propose THUNDER (Talking Heads Under Neural Differentiable Elocution Reconstruction), a 3D talking head avatar framework that introduces a novel supervision mechanism via differentiable sound production. First, we train a novel mesh-to-speech model that regresses audio from facial animation. Then, we incorporate this model into a diffusion-based talking avatar framework. During training, the mesh-to-speech model takes the generated animation and produces a sound that is compared to the input speech, creating a differentiable analysis-by-audio-synthesis supervision loop. Our extensive qualitative and quantitative experiments demonstrate that THUNDER significantly improves the quality of the lip-sync of talking head avatars while still allowing for generation of diverse, high-quality, expressive facial animations. The code and models will be available at https://thunder.is.tue.mpg.de/

📄 PDF Abstract BibTeX arXiv:2504.13386

Code (0)

등록된 구현이 없습니다.

Tasks

Audio Synthesis

Similar Papers 제목 키워드 기반

GaussianHeadTalk: Wobble-Free 3D Talking Heads with Audio Driven Gaussian Splatting

2025-12-11 · Madhav Agarwal, Mingtian Zhang, Laura Sevilla-Lara, Steven McDonagh arxiv

Speech-driven talking heads have recently emerged and enable interactive avatars. However, real-world applications are limited, as current methods achieve high visual fidelity but slow or fast yet temporally unstable. Di…

Image Generation

DEGAS: Detailed Expressions on Full-Body Gaussian Avatars

2024-08-20 · Zhijing Shao, Duotun Wang, Qing-Yao Tian, Yao-Dong Yang 외

Although neural rendering has made significant advances in creating lifelike, animatable full-body and head avatars, incorporating detailed expressions into full-body avatars remains largely unexplored. We present DEGAS,…

3DGSNeural Rendering

Comparative Analysis of Audio Feature Extraction for Real-Time Talking Portrait Synthesis

2024-11-20 · Pegah Salehi, Sajad Amouei Sheshkal, Vajira Thambawita, Sushant Gautam 외

This paper examines the integration of real-time talking-head generation for interviewer training, focusing on overcoming challenges in Audio Feature Extraction (AFE), which often introduces latency and limits responsive…

Talking Head Generation

Dual Audio-Centric Modality Coupling for Talking Head Generation

2025-03-26 · Ao Fu, Ziqi Ni, Yi Zhou

The generation of audio-driven talking head videos is a key challenge in computer vision and graphics, with applications in virtual avatars and digital media. Traditional approaches often struggle with capturing the comp…

NeRFTalking Head Generationtext-to-speechText to Speech

VASA-3D: Lifelike Audio-Driven Gaussian Head Avatars from a Single Image

2025-12-16 · Sicheng Xu, Guojun Chen, Jiaolong Yang, Yizhong Zhang 외 arxiv

We propose VASA-3D, an audio-driven, single-shot 3D head avatar generator. This research tackles two major challenges: capturing the subtle expression details present in real human faces, and reconstructing an intricate …