paper-with-me

홈 › Papers

Text2Lip: Progressive Lip-Synced Talking Face Generation from Text via Viseme-Guided Rendering

2025-08-04 · Xu Wang, Shengeng Tang, Fei Wang, Lechao Cheng, Dan Guo, Feng Xue, Richang Hong arxiv

Generating semantically coherent and visually accurate talking faces requires bridging the gap between linguistic meaning and facial articulation. Although audio-driven methods remain prevalent, their reliance on high-quality paired audio visual data and the inherent ambiguity in mapping acoustics to lip motion pose significant challenges in terms of scalability and robustness. To address these issues, we propose Text2Lip, a viseme-centric framework that constructs an interpretable phonetic-visual bridge by embedding textual input into structured viseme sequences. These mid-level units serve as a linguistically grounded prior for lip motion prediction. Furthermore, we design a progressive viseme-audio replacement strategy based on curriculum learning, enabling the model to gradually transition from real audio to pseudo-audio reconstructed from enhanced viseme features via cross-modal attention. This allows for robust generation in both audio-present and audio-free scenarios. Finally, a landmark-guided renderer synthesizes photorealistic facial videos with accurate lip synchronization. Extensive evaluations show that Text2Lip outperforms existing approaches in semantic fidelity, visual realism, and modality robustness, establishing a new paradigm for controllable and flexible talking face generation. Our project homepage is https://plyon1.github.io/Text2Lip/.

📄 PDF Abstract BibTeX arXiv:2508.02362

Code (0)

등록된 구현이 없습니다.

Tasks

Talking Face Generation

Similar Papers 제목 키워드 기반

Identity-Preserving Talking Face Generation with Landmark and Appearance Priors

2023-05-15 · CVPR 2023 1 · Weizhi Zhong, Chaowei Fang, Yinqi Cai, Pengxu Wei 외

Generating talking face videos from audio attracts lots of research interest. A few person-specific methods can generate vivid videos but require the target speaker's videos for training or fine-tuning. Existing person-g…

Face GenerationTalking Face Generation

Neural Text to Articulate Talk: Deep Text to Audiovisual Speech Synthesis achieving both Auditory and Photo-realism

2023-12-11 · Georgios Milis, Panagiotis P. Filntisis, Anastasios Roussos, Petros Maragos

Recent advances in deep learning for sequential data have given rise to fast and powerful models that produce realistic videos of talking humans. The state of the art in talking face generation focuses mainly on lip-sync…

Face GenerationLip ReadingSpeech SynthesisTalking Face Generation+2

High-fidelity and Lip-synced Talking Face Synthesis via Landmark-based Diffusion Model

2024-08-10 · Weizhi Zhong, Junfan Lin, Peixin Chen, Liang Lin 외

Audio-driven talking face video generation has attracted increasing attention due to its huge industrial potential. Some previous methods focus on learning a direct mapping from audio to visual content. Despite progress,…

Face GenerationTalking Face GenerationVideo Generation

Emotionally Enhanced Talking Face Generation

2023-03-21 · Sahil Goyal, Shagun Uppal, Sarthak Bhagat, Yi Yu 외

Several works have developed end-to-end pipelines for generating lip-synced talking faces with various real-world applications, such as teaching and language translation in videos. However, these prior works fail to crea…

Face GenerationTalking Face GenerationTalking Head Generation

Expressive Talking Head Generation With Granular Audio-Visual Control

2022-01-01 · CVPR 2022 1 · Borong Liang, Yan Pan, Zhizhi Guo, Hang Zhou 외

Generating expressive talking heads is essential for creating virtual humans. However, existing one- or few-shot methods focus on lip-sync and head motion, ignoring the emotional expressions that make talking faces r…

Talking Head Generation