paper-with-me

Papers

Learning Audio-Driven Viseme Dynamics for 3D Face Animation

2023-01-15 · Linchao Bao, Haoxian Zhang, Yue Qian, Tangli Xue, Changhai Chen, Xuefei Zhe, Di Kang

We present a novel audio-driven facial animation approach that can generate realistic lip-synchronized 3D facial animations from the input audio. Our approach learns viseme dynamics from speech videos, produces animator-friendly viseme curves, and supports multilingual speech inputs. The core of our approach is a novel parametric viseme fitting algorithm that utilizes phoneme priors to extract viseme parameters from speech videos. With the guidance of phonemes, the extracted viseme curves can better correlate with phonemes, thus more controllable and friendly to animators. To support multilingual speech inputs and generalizability to unseen voices, we take advantage of deep audio feature models pretrained on multiple languages to learn the mapping from audio to viseme curves. Our audio-to-curves mapping achieves state-of-the-art performance even when the input audio suffers from distortions of volume, pitch, speed, or noise. Lastly, a viseme scanning approach for acquiring high-fidelity viseme assets is presented for efficient speech animation production. We show that the predicted viseme curves can be applied to different viseme-rigged characters to yield various personalized animations with realistic and natural facial motions. Our approach is artist-friendly and can be easily integrated into typical animation production workflows including blendshape or bone based animation.

📄 PDF Abstract BibTeX arXiv:2301.06059

Code (0)

등록된 구현이 없습니다.

Tasks

3D Face Animation

Similar Papers 제목 키워드 기반

Learning Phonetic Context-Dependent Viseme for Enhancing Speech-Driven 3D Facial Animation

2025-07-28 · Hyung Kyu Kim, Hak Gu Kim arxiv

Speech-driven 3D facial animation aims to generate realistic facial movements synchronized with audio. Traditional methods primarily minimize reconstruction loss by aligning each frame with ground-truth. However, this fr…

3DiFACE: Synthesizing and Editing Holistic 3D Facial Animation

2025-09-30 · Balamurugan Thambiraja, Malte Prinzler, Sadegh Aliakbarian, Darren Cosker 외 arxiv

Creating personalized 3D animations with precise control and realistic head motions remains challenging for current speech-driven 3D facial animation methods. Editing these animations is especially complex and time consu…

A Bridge from Audio to Video: Phoneme-Viseme Alignment Allows Every Face to Speak Multiple Languages

2025-10-08 · Zibo Su, Kun Wei, Jiahua Li, Xu Yang 외 arxiv

Speech-driven talking face synthesis (TFS) focuses on generating lifelike facial animations from speech input. Current TFS models perform well in English but struggle with non-English languages, producing inaccurate mout…

Zero-shot Generalization

FACTS: Facial Animation Creation using the Transfer of Styles

2023-07-18 · Jack Saunders, Steven Caulkin, Vinay Namboodiri

The ability to accurately capture and express emotions is a critical aspect of creating believable characters in video games and other forms of entertainment. Traditionally, this animation has been achieved with artistic…

Text2Lip: Progressive Lip-Synced Talking Face Generation from Text via Viseme-Guided Rendering

2025-08-04 · Xu Wang, Shengeng Tang, Fei Wang, Lechao Cheng 외 arxiv

Generating semantically coherent and visually accurate talking faces requires bridging the gap between linguistic meaning and facial articulation. Although audio-driven methods remain prevalent, their reliance on high-qu…

Talking Face Generation