paper-with-me

홈 › Papers

Toward Fine-Grained Facial Control in 3D Talking Head Generation

2026-02-10 · Shaoyang Xie, Xiaofeng Cong, Baosheng Yu, Zhipeng Gui, Jie Gui, Yuan Yan Tang, James Tin-Yau Kwok arxiv

Audio-driven talking head generation is a core component of digital avatars, and 3D Gaussian Splatting has shown strong performance in real-time rendering of high-fidelity talking heads. However, achieving precise control over fine-grained facial movements remains a significant challenge, particularly due to lip-synchronization inaccuracies and facial jitter, both of which can contribute to the uncanny valley effect. To address these challenges, we propose Fine-Grained 3D Gaussian Splatting (FG-3DGS), a novel framework that enables temporally consistent and high-fidelity talking head generation. Our method introduces a frequency-aware disentanglement strategy to explicitly model facial regions based on their motion characteristics. Low-frequency regions, such as the cheeks, nose, and forehead, are jointly modeled using a standard MLP, while high-frequency regions, including the eyes and mouth, are captured separately using a dedicated network guided by facial area masks. The predicted motion dynamics, represented as Gaussian deltas, are applied to the static Gaussians to generate the final head frames, which are rendered via a rasterizer using frame-specific camera parameters. Additionally, a high-frequency-refined post-rendering alignment mechanism, learned from large-scale audio-video pairs by a pretrained model, is incorporated to enhance per-frame generation and achieve more accurate lip synchronization. Extensive experiments on widely used datasets for talking head generation demonstrate that our method outperforms recent state-of-the-art approaches in producing high-fidelity, lip-synced talking head videos.

📄 PDF Abstract BibTeX arXiv:2602.09736

Code (0)

등록된 구현이 없습니다.

Tasks

Talking Head Generation

Similar Papers 제목 키워드 기반

EmoDiffTalk:Emotion-aware Diffusion for Editable 3D Gaussian Talking Head

2025-11-30 · Chang Liu, Tianjiao Jing, Chengcheng Ma, Xuanqi Zhou 외 arxiv

Recent photo-realistic 3D talking head via 3D Gaussian Splatting still has significant shortcoming in emotional expression manipulation, especially for fine-grained and expansive dynamics emotional editing using multi-mo…

LES-Talker: Fine-Grained Emotion Editing for Talking Head Generation in Linear Emotion Space

2024-11-14 · Guanwen Feng, Zhihao Qian, Yunan Li, Siyu Jin 외

While existing one-shot talking head generation models have achieved progress in coarse-grained emotion editing, there is still a lack of fine-grained emotion editing models with high interpretability. We argue that for …

Talking Head Generation

Progressive Disentangled Representation Learning for Fine-Grained Controllable Talking Head Synthesis

2022-11-26 · CVPR 2023 1 · Duomin Wang, Yu Deng, Zixin Yin, Heung-Yeung Shum 외

We present a novel one-shot talking head synthesis method that achieves disentangled and fine-grained control over lip motion, eye gaze&blink, head pose, and emotional expression. We represent different motions via disen…

Contrastive LearningDisentanglementRepresentation Learning

Playmate: Flexible Control of Portrait Animation via 3D-Implicit Space Guided Diffusion

2025-02-11 · Xingpei Ma, Jiaran Cai, Yuansheng Guan, Shenneng Huang 외

Recent diffusion-based talking face generation models have demonstrated impressive potential in synthesizing videos that accurately match a speech audio clip with a given reference identity. However, existing approaches …

AttributeDisentanglementFace GenerationPortrait Animation+1

Talking Head Generation via AU-Guided Landmark Prediction

2025-09-24 · Shao-Yu Chang, Jingyi Xu, Hieu Le, Dimitris Samaras arxiv

We propose a two-stage framework for audio-driven talking head generation with fine-grained expression control via facial Action Units (AUs). Unlike prior methods relying on emotion labels or implicit AU conditioning, ou…

Talking Head Generation