paper-with-me

Papers

GSTalker: Real-time Audio-Driven Talking Face Generation via Deformable Gaussian Splatting

2024-04-29 · Bo Chen, Shoukang Hu, Qi Chen, Chenpeng Du, Ran Yi, Yanmin Qian, Xie Chen

We present GStalker, a 3D audio-driven talking face generation model with Gaussian Splatting for both fast training (40 minutes) and real-time rendering (125 FPS) with a 3$\sim$5 minute video for training material, in comparison with previous 2D and 3D NeRF-based modeling frameworks which require hours of training and seconds of rendering per frame. Specifically, GSTalker learns an audio-driven Gaussian deformation field to translate and transform 3D Gaussians to synchronize with audio information, in which multi-resolution hashing grid-based tri-plane and temporal smooth module are incorporated to learn accurate deformation for fine-grained facial details. In addition, a pose-conditioned deformation field is designed to model the stabilized torso. To enable efficient optimization of the condition Gaussian deformation field, we initialize 3D Gaussians by learning a coarse static Gaussian representation. Extensive experiments in person-specific videos with audio tracks validate that GSTalker can generate high-fidelity and audio-lips synchronized results with fast training and real-time rendering speed.

📄 PDF Abstract BibTeX arXiv:2404.19040

Code (0)

등록된 구현이 없습니다.

Tasks

Face GenerationNeRFTalking Face Generation

Similar Papers 제목 키워드 기반

EGSTalker: Real-Time Audio-Driven Talking Head Generation with Efficient Gaussian Deformation

2025-10-03 · Tianheng Zhu, Yinfeng Yu, Liejun Wang, Fuchun Sun 외 arxiv

This paper presents EGSTalker, a real-time audio-driven talking head generation framework based on 3D Gaussian Splatting (3DGS). Designed to enhance both speed and visual fidelity, EGSTalker requires only 3-5 minutes of …

Talking Head Generation

PGSTalker: Real-Time Audio-Driven Talking Head Generation via 3D Gaussian Splatting with Pixel-Aware Density Control

2025-09-21 · Tianheng Zhu, Yinfeng Yu, Liejun Wang, Fuchun Sun 외 arxiv

Audio-driven talking head generation is crucial for applications in virtual reality, digital avatars, and film production. While NeRF-based methods enable high-fidelity reconstruction, they suffer from low rendering effi…

Talking Head Generation

GaussianEmoTalker: Real-Time Emotional Talking Head Synthesis with Audio-Driven and Blendshape-Based 3D Gaussian Splatting

2026-07-01 · Haijie Yang, Zhenyu Zhang, Yixuan Dong, Jianjun Qian 외 arxiv

Audio-driven talking head synthesis has achieved impressive progress in lip synchronization and visual quality, yet generating expressive emotional avatars with controllable intensity remains challenging, especially unde…

Audio-Plane: Audio Factorization Plane Gaussian Splatting for Real-Time Talking Head Synthesis

2025-03-28 · Shuai Shen, Wanhua Li, Yunpeng Zhang, Weipeng Hu 외

Talking head synthesis has become a key research area in computer graphics and multimedia, yet most existing methods often struggle to balance generation quality with computational efficiency. In this paper, we present a…

Computational EfficiencyTalking Head Generation

FREAK: Frequency-modulated High-fidelity and Real-time Audio-driven Talking Portrait Synthesis

2025-03-06 · Ziqi Ni, Ao Fu, Yi Zhou

Achieving high-fidelity lip-speech synchronization in audio-driven talking portrait synthesis remains challenging. While multi-stage pipelines or diffusion models yield high-quality results, they suffer from high computa…

Audio-Visual Synchronization