paper-with-me

홈 › Papers

GaussianSpeech: Audio-Driven Gaussian Avatars

2024-11-27 · Shivangi Aneja, Artem Sevastopolsky, Tobias Kirschstein, Justus Thies, Angela Dai, Matthias Nießner

We introduce GaussianSpeech, a novel approach that synthesizes high-fidelity animation sequences of photo-realistic, personalized 3D human head avatars from spoken audio. To capture the expressive, detailed nature of human heads, including skin furrowing and finer-scale facial movements, we propose to couple speech signal with 3D Gaussian splatting to create realistic, temporally coherent motion sequences. We propose a compact and efficient 3DGS-based avatar representation that generates expression-dependent color and leverages wrinkle- and perceptually-based losses to synthesize facial details, including wrinkles that occur with different expressions. To enable sequence modeling of 3D Gaussian splats with audio, we devise an audio-conditioned transformer model capable of extracting lip and expression features directly from audio input. Due to the absence of high-quality datasets of talking humans in correspondence with audio, we captured a new large-scale multi-view dataset of audio-visual sequences of talking humans with native English accents and diverse facial geometry. GaussianSpeech consistently achieves state-of-the-art performance with visually natural motion at real time rendering rates, while encompassing diverse facial expressions and styles.

📄 PDF Abstract BibTeX arXiv:2411.18675

Code (1)

shivangi-aneja/GaussianSpeech 공식 구현 pytorch

Tasks

3DGS

Similar Papers 제목 키워드 기반

GaussianHeadTalk: Wobble-Free 3D Talking Heads with Audio Driven Gaussian Splatting

2025-12-11 · Madhav Agarwal, Mingtian Zhang, Laura Sevilla-Lara, Steven McDonagh arxiv

Speech-driven talking heads have recently emerged and enable interactive avatars. However, real-world applications are limited, as current methods achieve high visual fidelity but slow or fast yet temporally unstable. Di…

Image Generation

DEGAS: Detailed Expressions on Full-Body Gaussian Avatars

2024-08-20 · Zhijing Shao, Duotun Wang, Qing-Yao Tian, Yao-Dong Yang 외

Although neural rendering has made significant advances in creating lifelike, animatable full-body and head avatars, incorporating detailed expressions into full-body avatars remains largely unexplored. We present DEGAS,…

3DGSNeural Rendering

PGSTalker: Real-Time Audio-Driven Talking Head Generation via 3D Gaussian Splatting with Pixel-Aware Density Control

2025-09-21 · Tianheng Zhu, Yinfeng Yu, Liejun Wang, Fuchun Sun 외 arxiv

Audio-driven talking head generation is crucial for applications in virtual reality, digital avatars, and film production. While NeRF-based methods enable high-fidelity reconstruction, they suffer from low rendering effi…

Talking Head Generation

READ Avatars: Realistic Emotion-controllable Audio Driven Avatars

2023-03-01 · Jack Saunders, Vinay Namboodiri

We present READ Avatars, a 3D-based approach for generating 2D avatars that are driven by audio input with direct and granular control over the emotion. Previous methods are unable to achieve realistic animation due to t…

Audio2Face-3D: Audio-driven Realistic Facial Animation For Digital Avatars

2025-08-22 · NVIDIA, :, Chaeyeon Chung, Ilya Fedorov 외 arxiv

Audio-driven facial animation presents an effective solution for animating digital avatars. In this paper, we detail the technical aspects of NVIDIA Audio2Face-3D, including data acquisition, network architecture, retarg…