paper-with-me

Papers

SyncTalk: The Devil is in the Synchronization for Talking Head Synthesis

2023-11-29 · CVPR 2024 1 · Ziqiao Peng, Wentao Hu, Yue Shi, Xiangyu Zhu, Xiaomei Zhang, Hao Zhao, Jun He, Hongyan Liu, Zhaoxin Fan

Achieving high synchronization in the synthesis of realistic, speech-driven talking head videos presents a significant challenge. Traditional Generative Adversarial Networks (GAN) struggle to maintain consistent facial identity, while Neural Radiance Fields (NeRF) methods, although they can address this issue, often produce mismatched lip movements, inadequate facial expressions, and unstable head poses. A lifelike talking head requires synchronized coordination of subject identity, lip movements, facial expressions, and head poses. The absence of these synchronizations is a fundamental flaw, leading to unrealistic and artificial outcomes. To address the critical issue of synchronization, identified as the "devil" in creating realistic talking heads, we introduce SyncTalk. This NeRF-based method effectively maintains subject identity, enhancing synchronization and realism in talking head synthesis. SyncTalk employs a Face-Sync Controller to align lip movements with speech and innovatively uses a 3D facial blendshape model to capture accurate facial expressions. Our Head-Sync Stabilizer optimizes head poses, achieving more natural head movements. The Portrait-Sync Generator restores hair details and blends the generated head with the torso for a seamless visual experience. Extensive experiments and user studies demonstrate that SyncTalk outperforms state-of-the-art methods in synchronization and realism. We recommend watching the supplementary video: https://ziqiaopeng.github.io/synctalk

📄 PDF Abstract BibTeX arXiv:2311.17590

Code (1)

ZiqiaoPeng/SyncTalk 공식 구현 pytorch

Tasks

NeRFTalking Face GenerationTalking Head Generation

Methods 이 논문이 사용한 방법론

ALIGN In the ALIGN method, visual and language representations are jointly trained from noisy image alt-text data. The image and text encoders are learned via contrastive loss…

Similar Papers 제목 키워드 기반

SyncTalk++: High-Fidelity and Efficient Synchronized Talking Heads Synthesis Using Gaussian Splatting

2025-06-17 · Ziqiao Peng, Wentao Hu, Junyuan Ma, Xiangyu Zhu 외

Achieving high synchronization in the synthesis of realistic, speech-driven talking head videos presents a significant challenge. A lifelike talking head requires synchronized coordination of subject identity, lip moveme…

EmoTalkingGaussian: Continuous Emotion-conditioned Talking Head Synthesis

2025-02-02 · Junuk Cha, Seongro Yoon, Valeriya Strizhkova, Francois Bremond 외

3D Gaussian splatting-based talking head synthesis has recently gained attention for its ability to render high-fidelity images with real-time inference speed. However, since it is typically trained on only a short video…

Self-Supervised LearningSSIMtext-to-speechText to Speech

PointTalk: Audio-Driven Dynamic Lip Point Cloud for 3D Gaussian-based Talking Head Synthesis

2024-12-11 · Yifan Xie, Tao Feng, Xin Zhang, Xiangyang Luo 외

Talking head synthesis with arbitrary speech audio is a crucial challenge in the field of digital humans. Recently, methods based on radiance fields have received increasing attention due to their ability to synthesize h…

SyncTalkFace: Talking Face Generation with Precise Lip-Syncing via Audio-Lip Memory

2022-11-02 · Se Jin Park, Minsu Kim, Joanna Hong, Jeongsoo Choi 외

The challenge of talking face generation from speech lies in aligning two different modal information, audio and video, such that the mouth region corresponds to input audio. Previous methods either exploit audio-visual …

Audio-Visual SynchronizationFace GenerationRepresentation LearningTalking Face Generation

GaussianEmoTalker: Real-Time Emotional Talking Head Synthesis with Audio-Driven and Blendshape-Based 3D Gaussian Splatting

2026-07-01 · Haijie Yang, Zhenyu Zhang, Yixuan Dong, Jianjun Qian 외 arxiv

Audio-driven talking head synthesis has achieved impressive progress in lip synchronization and visual quality, yet generating expressive emotional avatars with controllable intensity remains challenging, especially unde…