paper-with-me

홈 › Papers

Realistic Lip Motion Generation Based on 3D Dynamic Viseme and Coarticulation Modeling for Human-Robot Interaction

2026-04-02 · Sheng Li, Jingcheng Huang, Min Li arxiv

Realistic lip synchronization is essential for the natural human-robot non-verbal interaction of humanoid robots. Motivated by this need, this paper presents a lip motion generation framework based on 3D dynamic viseme and coarticulation modeling. By analyzing Chinese pronunciation theory, a 3D dynamic viseme library is constructed based on the ARKit standard, which offers coherent prior trajectories of lips. To resolve motion conflicts within continuous speech streams, a coarticulation mechanism is developed by incorporating initial-final (Shengmu-Yunmu) decoupling and energy modulation. After developing a strategy to retarget high-dimensional spatial lip motion to a 14-DOF lip actuation system of a humanoid head platform, the efficiency and accuracy of the proposed architecture is experimentally validated and demonstrated with quantitative ablation experiments using the metrics of the Pearson Correlation Coefficient (PCC) and the Mean Absolute Jerk (MAJ). This research offers a lightweight, efficient, and highly practical paradigm for the speech-driven lip motion generation of humanoid robots. The 3D dynamic viseme library and real-world deployment videos are available at {https://github.com/yuesheng21/Phoneme-to-Lip-14DOF}

📄 PDF Abstract BibTeX arXiv:2604.01756

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Learning Phonetic Context-Dependent Viseme for Enhancing Speech-Driven 3D Facial Animation

2025-07-28 · Hyung Kyu Kim, Hak Gu Kim arxiv

Speech-driven 3D facial animation aims to generate realistic facial movements synchronized with audio. Traditional methods primarily minimize reconstruction loss by aligning each frame with ground-truth. However, this fr…

Learning Audio-Driven Viseme Dynamics for 3D Face Animation

2023-01-15 · Linchao Bao, Haoxian Zhang, Yue Qian, Tangli Xue 외

We present a novel audio-driven facial animation approach that can generate realistic lip-synchronized 3D facial animations from the input audio. Our approach learns viseme dynamics from speech videos, produces animator-…

3D Face Animation

Text2Lip: Progressive Lip-Synced Talking Face Generation from Text via Viseme-Guided Rendering

2025-08-04 · Xu Wang, Shengeng Tang, Fei Wang, Lechao Cheng 외 arxiv

Generating semantically coherent and visually accurate talking faces requires bridging the gap between linguistic meaning and facial articulation. Although audio-driven methods remain prevalent, their reliance on high-qu…

Talking Face Generation

VedicTHG: Symbolic Vedic Computation for Low-Resource Talking-Head Generation in Educational Avatars

2026-02-09 · Vineet Kumar Rakesh, Ahana Bhattacharjee, Soumya Mazumdar, Tapas Samanta 외 arxiv

Talking-head avatars are increasingly adopted in educational technology to deliver content with social presence and improved engagement. However, many recent talking-head generation (THG) methods rely on GPU-centric neur…

Towards a Quantitative Analysis of Coarticulation with a Phoneme-to-Articulatory Model

2024-08-10 · Chaofei Fan, Jaimie M. Henderson, Chris Manning, Francis R. Willett

Prior coarticulation studies focus mainly on limited phonemic sequences and specific articulators, providing only approximate descriptions of the temporal extent and magnitude of coarticulation. This paper is an initial …