paper-with-me

홈 › Papers

LipFormer: High-Fidelity and Generalizable Talking Face Generation With a Pre-Learned Facial Codebook

2023-01-01 · CVPR 2023 1 · Jiayu Wang, Kang Zhao, Shiwei Zhang, Yingya Zhang, Yujun Shen, Deli Zhao, Jingren Zhou

Generating a talking face video from the input audio sequence is a practical yet challenging task. Most existing methods either fail to capture fine facial details or need to train a specific model for each identity. We argue that a codebook pre-learned on high-quality face images can serve as a useful prior that facilitates high-fidelity and generalizable talking head synthesis. Thanks to the strong capability of the codebook in representing face textures, we simplify the talking face generation task as finding proper lip-codes to characterize the variation of lips during a portrait talking. To this end, we propose LipFormer, a transformer-based framework, to model the audio-visual coherence and predict the lip-codes sequence based on the input audio features. We further introduce an adaptive face warping module, which helps warp the reference face to the target pose in the feature space, to alleviate the difficulty of lip-code prediction under different poses. By this means, LipFormer can make better use of the pre-learned priors in images and is robust to posture change. Extensive experiments show that LipFormer can produce more realistic talking face videos compared to previous methods and faithfully generalize to unseen identities.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Tasks

Face GenerationTalking Face Generation

Methods 이 논문이 사용한 방법론

fail 설명 없음

Similar Papers 제목 키워드 기반

Text-based Talking Video Editing with Cascaded Conditional Diffusion

2024-07-20 · Bo Han, Heqing Zou, Haoyang Li, Guangcong Wang 외

Text-based talking-head video editing aims to efficiently insert, delete, and substitute segments of talking videos through a user-friendly text editing approach. It is challenging because of \textbf{1)} generalizable ta…

Video Editing

G4G:A Generic Framework for High Fidelity Talking Face Generation with Fine-grained Intra-modal Alignment

2024-02-28 · Juan Zhang, Jiahao Chen, Cheng Wang, Zhiwang Yu 외

Despite numerous completed studies, achieving high fidelity talking face generation with highly synchronized lip movements corresponding to arbitrary audio remains a significant challenge in the field. The shortcomings o…

Face GenerationTalking Face Generation

GeneFace: Generalized and High-Fidelity Audio-Driven 3D Talking Face Synthesis

2023-01-31 · Zhenhui Ye, Ziyue Jiang, Yi Ren, Jinglin Liu 외

Generating photo-realistic video portrait with arbitrary speech audio is a crucial problem in film-making and virtual reality. Recently, several works explore the usage of neural radiance field in this task to improve 3D…

Face GenerationLip ReadingNeRFTalking Face Generation+1

ClipFormer: Key-Value Clipping of Transformers on Memristive Crossbars for Write Noise Mitigation

2024-02-04 · Abhiroop Bhattacharjee, Abhishek Moitra, Priyadarshini Panda

Transformers have revolutionized various real-world applications from natural language processing to computer vision. However, traditional von-Neumann computing paradigm faces memory and bandwidth limitations in accelera…

Audio-Driven Talking Face Generation with Blink Embedding and Hash Grid Landmarks Encoding

2026-01-26 · Yuhui Zhang, Hui Yu, Wei Liang, Sunjie Zhang arxiv

Dynamic Neural Radiance Fields (NeRF) have demonstrated considerable success in generating high-fidelity 3D models of talking portraits. Despite significant advancements in the rendering speed and generation quality, cha…

Talking Face Generation