paper-with-me

Papers

Learn2Talk: 3D Talking Face Learns from 2D Talking Face

2024-04-19 · Yixiang Zhuang, Baoping Cheng, Yao Cheng, Yuntao Jin, Renshuai Liu, Chengyang Li, Xuan Cheng, Jing Liao, Juncong Lin

Speech-driven facial animation methods usually contain two main classes, 3D and 2D talking face, both of which attract considerable research attention in recent years. However, to the best of our knowledge, the research on 3D talking face does not go deeper as 2D talking face, in the aspect of lip-synchronization (lip-sync) and speech perception. To mind the gap between the two sub-fields, we propose a learning framework named Learn2Talk, which can construct a better 3D talking face network by exploiting two expertise points from the field of 2D talking face. Firstly, inspired by the audio-video sync network, a 3D sync-lip expert model is devised for the pursuit of lip-sync between audio and 3D facial motion. Secondly, a teacher model selected from 2D talking face methods is used to guide the training of the audio-to-3D motions regression network to yield more 3D vertex accuracy. Extensive experiments show the advantages of the proposed framework in terms of lip-sync, vertex accuracy and speech perception, compared with state-of-the-arts. Finally, we show two applications of the proposed framework: audio-visual speech recognition and speech-driven 3D Gaussian Splatting based avatar animation.

📄 PDF Abstract BibTeX arXiv:2404.12888

Code (0)

등록된 구현이 없습니다.

Tasks

Audio-Visual Speech Recognitionspeech-recognitionSpeech RecognitionVisual Speech Recognition

Similar Papers 제목 키워드 기반

FacEDiT: Unified Talking Face Editing and Generation via Facial Motion Infilling

2025-12-16 · Kim Sung-Bin, Joohyun Chang, David Harwath, Tae-Hyun Oh arxiv

Talking face editing and face generation have often been studied as distinct problems. In this work, we propose viewing both not as separate tasks but as subtasks of a unifying formulation, speech-conditional facial moti…

Talking Face Generation

Imitating Arbitrary Talking Style for Realistic Audio-DrivenTalking Face Synthesis

2021-10-30 · Haozhe Wu, Jia Jia, Haoyu Wang, Yishun Dou 외

People talk with diversified styles. For one piece of speech, different talking styles exhibit significant differences in the facial and head pose movements. For example, the "excited" style usually talks with the mouth …

Face Generation

Expressive Talking Head Generation With Granular Audio-Visual Control

2022-01-01 · CVPR 2022 1 · Borong Liang, Yan Pan, Zhizhi Guo, Hang Zhou 외

Generating expressive talking heads is essential for creating virtual humans. However, existing one- or few-shot methods focus on lip-sync and head motion, ignoring the emotional expressions that make talking faces r…

Talking Head Generation

AnyoneNet: Synchronized Speech and Talking Head Generation for Arbitrary Person

2021-08-09 · Xinsheng Wang, Qicong Xie, Jihua Zhu, Lei Xie 외

Automatically generating videos in which synthesized speech is synchronized with lip movements in a talking head has great potential in many human-computer interaction scenarios. In this paper, we present an automatic me…

Talking Head Generationtext-to-speechText to Speech

AVI-Talking: Learning Audio-Visual Instructions for Expressive 3D Talking Face Generation

2024-02-25 · Yasheng Sun, Wenqing Chu, Hang Zhou, Kaisiyuan Wang 외

While considerable progress has been made in achieving accurate lip synchronization for 3D speech-driven talking face generation, the task of incorporating expressive facial detail synthesis aligned with the speaker's sp…

Face GenerationHallucinationTalking Face Generation