paper-with-me

Papers

LipSync3D: Data-Efficient Learning of Personalized 3D Talking Faces from Video using Pose and Lighting Normalization

2021-06-08 · CVPR 2021 1 · Avisek Lahiri, Vivek Kwatra, Christian Frueh, John Lewis, Chris Bregler

In this paper, we present a video-based learning framework for animating personalized 3D talking faces from audio. We introduce two training-time data normalizations that significantly improve data sample efficiency. First, we isolate and represent faces in a normalized space that decouples 3D geometry, head pose, and texture. This decomposes the prediction problem into regressions over the 3D face shape and the corresponding 2D texture atlas. Second, we leverage facial symmetry and approximate albedo constancy of skin to isolate and remove spatio-temporal lighting variations. Together, these normalizations allow simple networks to generate high fidelity lip-sync videos under novel ambient illumination while training with just a single speaker-specific video. Further, to stabilize temporal dynamics, we introduce an auto-regressive approach that conditions the model on its previous visual state. Human ratings and objective metrics demonstrate that our method outperforms contemporary state-of-the-art audio-driven video reenactment benchmarks in terms of realism, lip-sync and visual quality scores. We illustrate several applications enabled by our framework.

📄 PDF Abstract BibTeX arXiv:2106.04185

Code (0)

등록된 구현이 없습니다.

Tasks

3D geometry

Similar Papers 제목 키워드 기반

Audio-driven Talking Face Video Generation with Learning-based Personalized Head Pose

2020-02-24 · Ran Yi, Zipeng Ye, Juyong Zhang, Hujun Bao 외

Real-world talking faces often accompany with natural head movement. However, most existing talking face video generation methods only consider facial animation with fixed head pose. In this paper, we address this proble…

3D Face AnimationVideo Generation

Livatar-1: Real-Time Talking Heads Generation with Tailored Flow Matching

2025-07-22 · Haiyang Liu, Xiaolin Hong, Xuancheng Yang, Yudi Ruan 외 arxiv

We present Livatar, a real-time audio-driven talking heads videos generation framework. Existing baselines suffer from limited lip-sync accuracy and long-term pose drift. We address these limitations with a flow matching…

StyleLipSync: Style-based Personalized Lip-sync Video Generation

2023-04-30 · ICCV 2023 1 · Taekyung Ki, Dongchan Min

In this paper, we present StyleLipSync, a style-based personalized lip-sync video generative model that can generate identity-agnostic lip-synchronizing video from arbitrary audio. To generate a video of arbitrary identi…

Video Generation

A Lip Sync Expert Is All You Need for Speech to Lip Generation In The Wild

2020-08-23 · K R Prajwal, Rudrabha Mukhopadhyay, Vinay Namboodiri, C. V. Jawahar

In this work, we investigate the problem of lip-syncing a talking face video of an arbitrary identity to match a target speech segment. Current works excel at producing accurate lip movements on a static image or videos …

AllMORPHTalking Face GenerationTalking Head Generation+1

MyPortrait: Morphable Prior-Guided Personalized Portrait Generation

2023-12-05 · Bo Ding, Zhenfeng Fan, Shuang Yang, Shihong Xia

Generating realistic talking faces is an interesting and long-standing topic in the field of computer vision. Although significant progress has been made, it is still challenging to generate high-quality dynamic faces wi…