paper-with-me

Papers

EDTalk: Efficient Disentanglement for Emotional Talking Head Synthesis

2024-04-02 · Shuai Tan, Bin Ji, Mengxiao Bi, Ye Pan

Achieving disentangled control over multiple facial motions and accommodating diverse input modalities greatly enhances the application and entertainment of the talking head generation. This necessitates a deep exploration of the decoupling space for facial features, ensuring that they a) operate independently without mutual interference and b) can be preserved to share with different modal input, both aspects often neglected in existing methods. To address this gap, this paper proposes a novel Efficient Disentanglement framework for Talking head generation (EDTalk). Our framework enables individual manipulation of mouth shape, head pose, and emotional expression, conditioned on video or audio inputs. Specifically, we employ three lightweight modules to decompose the facial dynamics into three distinct latent spaces representing mouth, pose, and expression, respectively. Each space is characterized by a set of learnable bases whose linear combinations define specific motions. To ensure independence and accelerate training, we enforce orthogonality among bases and devise an efficient training strategy to allocate motion responsibilities to each space without relying on external knowledge. The learned bases are then stored in corresponding banks, enabling shared visual priors with audio input. Furthermore, considering the properties of each space, we propose an Audio-to-Motion module for audio-driven talking head synthesis. Experiments are conducted to demonstrate the effectiveness of EDTalk. We recommend watching the project website: https://tanshuai0219.github.io/EDTalk/

📄 PDF Abstract BibTeX arXiv:2404.01647

Code (0)

등록된 구현이 없습니다.

Tasks

DisentanglementTalking Head Generation

Methods 이 논문이 사용한 방법론

SET Dynamic Sparse Training method where weight mask is updated randomly periodically

Similar Papers 제목 키워드 기반

EDTalk++: Full Disentanglement for Controllable Talking Head Synthesis

2025-08-19 · Shuai Tan, Bin Ji arxiv

Achieving disentangled control over multiple facial motions and accommodating diverse input modalities greatly enhances the application and entertainment of the talking head generation. This necessitates a deep explorati…

Talking Head Generation

EmbedTalk: Triplane-Free Talking Head Synthesis using Embedding-Driven Gaussian Deformation

2026-03-08 · Arpita Saggar, Jonathan C. Darling, Duygu Sarikaya, David C. Hogg arxiv

Real-time talking head synthesis increasingly relies on deformable 3D Gaussian Splatting (3DGS) due to its low latency. Tri-planes are the standard choice for encoding Gaussians prior to deformation, since they provide a…

SEDTalker: Emotion-Aware 3D Facial Animation Using Frame-Level Speech Emotion Diarization

2026-04-14 · Farzaneh Jafari, Stefano Berretti, Anup Basu arxiv

We introduce SEDTalker, an emotion-aware framework for speech-driven 3D facial animation that leverages frame-level speech emotion diarization to achieve fine-grained expressive control. Unlike prior approaches that rely…

Talking Head GenerationEmotion Recognition

Uncertainty-Aware 3D Emotional Talking Face Synthesis with Emotion Prior Distillation

2026-01-27 · Nanhan Shen, Zhilei Liu arxiv

Emotional Talking Face synthesis is pivotal in multimedia and signal processing, yet existing 3D methods suffer from two critical challenges: poor audio-vision emotion alignment, manifested as difficult audio emotion ext…

Progressive Disentangled Representation Learning for Fine-Grained Controllable Talking Head Synthesis

2022-11-26 · CVPR 2023 1 · Duomin Wang, Yu Deng, Zixin Yin, Heung-Yeung Shum 외

We present a novel one-shot talking head synthesis method that achieves disentangled and fine-grained control over lip motion, eye gaze&blink, head pose, and emotional expression. We represent different motions via disen…

Contrastive LearningDisentanglementRepresentation Learning