paper-with-me

Papers

Media2Face: Co-speech Facial Animation Generation With Multi-Modality Guidance

2024-01-28 · Qingcheng Zhao, Pengyu Long, Qixuan Zhang, Dafei Qin, Han Liang, Longwen Zhang, Yingliang Zhang, Jingyi Yu, Lan Xu

The synthesis of 3D facial animations from speech has garnered considerable attention. Due to the scarcity of high-quality 4D facial data and well-annotated abundant multi-modality labels, previous methods often suffer from limited realism and a lack of lexible conditioning. We address this challenge through a trilogy. We first introduce Generalized Neural Parametric Facial Asset (GNPFA), an efficient variational auto-encoder mapping facial geometry and images to a highly generalized expression latent space, decoupling expressions and identities. Then, we utilize GNPFA to extract high-quality expressions and accurate head poses from a large array of videos. This presents the M2F-D dataset, a large, diverse, and scan-level co-speech 3D facial animation dataset with well-annotated emotional and style labels. Finally, we propose Media2Face, a diffusion model in GNPFA latent space for co-speech facial animation generation, accepting rich multi-modality guidances from audio, text, and image. Extensive experiments demonstrate that our model not only achieves high fidelity in facial animation synthesis but also broadens the scope of expressiveness and style adaptability in 3D facial animation.

📄 PDF Abstract BibTeX arXiv:2401.15687

Code (0)

등록된 구현이 없습니다.

Methods 이 논문이 사용한 방법론

Diffusion Diffusion models generate samples by gradually removing noise from a signal, and their training objective can be expressed as a reweighted variational lower-bound…

Similar Papers 제목 키워드 기반

AudioFace: Language-Assisted Speech-Driven Facial Animation with Multimodal Language Models

2026-05-08 · Kai Zheng, Zejian Kang, Rui Mao, Hongyuan Zou 외 arxiv

Speech-driven facial animation requires accurate correspondence between acoustic signals and facial motion, especially for articulation-related mouth movements. However, directly mapping speech audio to facial coefficien…

EmoDiffusion: Enhancing Emotional 3D Facial Animation with Latent Diffusion Models

2025-03-14 · Yixuan Zhang, Qing Chang, Yuxi Wang, Guang Chen 외

Speech-driven 3D facial animation seeks to produce lifelike facial expressions that are synchronized with the speech content and its emotional nuances, finding applications in various multimedia fields. However, previous…

DiffusionTalker: Personalization and Acceleration for Speech-Driven 3D Face Diffuser

2023-11-28 · Peng Chen, Xiaobao Wei, Ming Lu, Yitong Zhu 외

Speech-driven 3D facial animation has been an attractive task in both academia and industry. Traditional methods mostly focus on learning a deterministic mapping from speech to animation. Recent approaches start to consi…

3D Face AnimationContrastive LearningKnowledge Distillation

Identity-Preserving Realistic Talking Face Generation

2020-05-25 · Sanjana Sinha, Sandika Biswas, Brojeshwar Bhowmick

Speech-driven facial animation is useful for a variety of applications such as telepresence, chatbots, etc. The necessary attributes of having a realistic face animation are 1) audio-visual synchronization (2) identity p…

Audio-Visual SynchronizationFace GenerationImage ReconstructionTalking Face Generation

DF-3DFace: One-to-Many Speech Synchronized 3D Face Animation with Diffusion

2023-08-23 · Se Jin Park, Joanna Hong, Minsu Kim, Yong Man Ro

Speech-driven 3D facial animation has gained significant attention for its ability to create realistic and expressive facial animations in 3D space based on speech. Learning-based methods have shown promising progress in…

3D Face Animation