paper-with-me

홈 › Papers

G4G:A Generic Framework for High Fidelity Talking Face Generation with Fine-grained Intra-modal Alignment

2024-02-28 · Juan Zhang, Jiahao Chen, Cheng Wang, Zhiwang Yu, Tangquan Qi, Di wu

Despite numerous completed studies, achieving high fidelity talking face generation with highly synchronized lip movements corresponding to arbitrary audio remains a significant challenge in the field. The shortcomings of published studies continue to confuse many researchers. This paper introduces G4G, a generic framework for high fidelity talking face generation with fine-grained intra-modal alignment. G4G can reenact the high fidelity of original video while producing highly synchronized lip movements regardless of given audio tones or volumes. The key to G4G's success is the use of a diagonal matrix to enhance the ordinary alignment of audio-image intra-modal features, which significantly increases the comparative learning between positive and negative samples. Additionally, a multi-scaled supervision module is introduced to comprehensively reenact the perceptional fidelity of original video across the facial region while emphasizing the synchronization of lip movements and the input audio. A fusion network is then used to further fuse the facial region and the rest. Our experimental results demonstrate significant achievements in reenactment of original video quality as well as highly synchronized talking lips. G4G is an outperforming generic framework that can produce talking videos competitively closer to ground truth level than current state-of-the-art methods.

📄 PDF Abstract BibTeX arXiv:2402.18122

Code (0)

등록된 구현이 없습니다.

Tasks

Face GenerationTalking Face Generation

Similar Papers 제목 키워드 기반

Ada-TTA: Towards Adaptive High-Quality Text-to-Talking Avatar Synthesis

2023-06-06 · Zhenhui Ye, Ziyue Jiang, Yi Ren, Jinglin Liu 외

We are interested in a novel task, namely low-resource text-to-talking avatar. Given only a few-minute-long talking person video with the audio track as the training data and arbitrary texts as the driving input, we aim …

Neural Renderingtext-to-speechText to SpeechVideo Generation+1

LipFormer: High-Fidelity and Generalizable Talking Face Generation With a Pre-Learned Facial Codebook

2023-01-01 · CVPR 2023 1 · Jiayu Wang, Kang Zhao, Shiwei Zhang, Yingya Zhang 외

Generating a talking face video from the input audio sequence is a practical yet challenging task. Most existing methods either fail to capture fine facial details or need to train a specific model for each identity.…

Face GenerationTalking Face Generation

ScanTalk: 3D Talking Heads from Unregistered Scans

2024-03-16 · Federico Nocentini, Thomas Besnier, Claudio Ferrari, Sylvain Arguillere 외

Speech-driven 3D talking heads generation has emerged as a significant area of interest among researchers, presenting numerous challenges. Existing methods are constrained by animating faces with fixed topologies, wherei…

GeneFace: Generalized and High-Fidelity Audio-Driven 3D Talking Face Synthesis

2023-01-31 · Zhenhui Ye, Ziyue Jiang, Yi Ren, Jinglin Liu 외

Generating photo-realistic video portrait with arbitrary speech audio is a crucial problem in film-making and virtual reality. Recently, several works explore the usage of neural radiance field in this task to improve 3D…

Face GenerationLip ReadingNeRFTalking Face Generation+1

Lightweight High-Fidelity Low-Bitrate Talking Face Compression for 3D Video Conference

2026-01-29 · Jianglong Li, Jun Xu, Bingcong Lu, Zhengxue Cheng 외 arxiv

The demand for immersive and interactive communication has driven advancements in 3D video conferencing, yet achieving high-fidelity 3D talking face representation at low bitrates remains a challenge. Traditional 2D vide…