paper-with-me

홈 › Papers

RealTalk: Realistic Emotion-Aware Lifelike Talking-Head Synthesis

2025-08-16 · Wenqing Wang, Yun Fu arxiv

Emotion is a critical component of artificial social intelligence. However, while current methods excel in lip synchronization and image quality, they often fail to generate accurate and controllable emotional expressions while preserving the subject's identity. To address this challenge, we introduce RealTalk, a novel framework for synthesizing emotional talking heads with high emotion accuracy, enhanced emotion controllability, and robust identity preservation. RealTalk employs a variational autoencoder (VAE) to generate 3D facial landmarks from driving audio, which are concatenated with emotion-label embeddings using a ResNet-based landmark deformation model (LDM) to produce emotional landmarks. These landmarks and facial blendshape coefficients jointly condition a novel tri-plane attention Neural Radiance Field (NeRF) to synthesize highly realistic emotional talking heads. Extensive experiments demonstrate that RealTalk outperforms existing methods in emotion accuracy, controllability, and identity preservation, advancing the development of socially intelligent AI systems.

📄 PDF Abstract BibTeX arXiv:2508.12163

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

EAMM: One-Shot Emotional Talking Face via Audio-Based Emotion-Aware Motion Model

2022-05-30 · Xinya Ji, Hang Zhou, Kaisiyuan Wang, Qianyi Wu 외

Although significant progress has been made to audio-driven talking face generation, existing methods either neglect facial emotion or cannot be applied to arbitrary subjects. In this paper, we propose the Emotion-Aware …

Face GenerationTalking Face Generation

FlowVQTalker: High-Quality Emotional Talking Face Generation through Normalizing Flow and Quantization

2024-03-11 · CVPR 2024 1 · Shuai Tan, Bin Ji, Ye Pan

Generating emotional talking faces is a practical yet challenging endeavor. To create a lifelike avatar, we draw upon two critical insights from a human perspective: 1) The connection between audio and the non-determinis…

Face GenerationQuantizationTalking Face Generation

EmoDiffusion: Enhancing Emotional 3D Facial Animation with Latent Diffusion Models

2025-03-14 · Yixuan Zhang, Qing Chang, Yuxi Wang, Guang Chen 외

Speech-driven 3D facial animation seeks to produce lifelike facial expressions that are synchronized with the speech content and its emotional nuances, finding applications in various multimedia fields. However, previous…

EmoDiffTalk:Emotion-aware Diffusion for Editable 3D Gaussian Talking Head

2025-11-30 · Chang Liu, Tianjiao Jing, Chengcheng Ma, Xuanqi Zhou 외 arxiv

Recent photo-realistic 3D talking head via 3D Gaussian Splatting still has significant shortcoming in emotional expression manipulation, especially for fine-grained and expansive dynamics emotional editing using multi-mo…

MEMO: Memory-Guided Diffusion for Expressive Talking Video Generation

2024-12-05 · Longtao Zheng, Yifan Zhang, Hanzhong Guo, Jiachun Pan 외

Recent advances in video diffusion models have unlocked new potential for realistic audio-driven talking video generation. However, achieving seamless audio-lip synchronization, maintaining long-term identity consistency…

Portrait AnimationVideo Generation