paper-with-me

Papers

Parametric Implicit Face Representation for Audio-Driven Facial Reenactment

2023-06-13 · CVPR 2023 1 · Ricong Huang, Peiwen Lai, Yipeng Qin, Guanbin Li

Audio-driven facial reenactment is a crucial technique that has a range of applications in film-making, virtual avatars and video conferences. Existing works either employ explicit intermediate face representations (e.g., 2D facial landmarks or 3D face models) or implicit ones (e.g., Neural Radiance Fields), thus suffering from the trade-offs between interpretability and expressive power, hence between controllability and quality of the results. In this work, we break these trade-offs with our novel parametric implicit face representation and propose a novel audio-driven facial reenactment framework that is both controllable and can generate high-quality talking heads. Specifically, our parametric implicit representation parameterizes the implicit representation with interpretable parameters of 3D face models, thereby taking the best of both explicit and implicit methods. In addition, we propose several new techniques to improve the three components of our framework, including i) incorporating contextual information into the audio-to-expression parameters encoding; ii) using conditional image synthesis to parameterize the implicit representation and implementing it with an innovative tri-plane structure for efficient learning; iii) formulating facial reenactment as a conditional image inpainting problem and proposing a novel data augmentation technique to improve model generalizability. Extensive experiments demonstrate that our method can generate more realistic results than previous methods with greater fidelity to the identities and talking styles of speakers.

📄 PDF Abstract BibTeX arXiv:2306.07579

Code (0)

등록된 구현이 없습니다.

Tasks

Data AugmentationImage GenerationImage Inpainting

Methods 이 논문이 사용한 방법론

Inpainting Train a convolutional neural network to generate the contents of an arbitrary image region conditioned on its surroundings.

Similar Papers 제목 키워드 기반

EAMM: One-Shot Emotional Talking Face via Audio-Based Emotion-Aware Motion Model

2022-05-30 · Xinya Ji, Hang Zhou, Kaisiyuan Wang, Qianyi Wu 외

Although significant progress has been made to audio-driven talking face generation, existing methods either neglect facial emotion or cannot be applied to arbitrary subjects. In this paper, we propose the Emotion-Aware …

Face GenerationTalking Face Generation

FaceTalk: Audio-Driven Motion Diffusion for Neural Parametric Head Models

2023-12-13 · CVPR 2024 1 · Shivangi Aneja, Justus Thies, Angela Dai, Matthias Nießner

We introduce FaceTalk, a novel generative approach designed for synthesizing high-fidelity 3D motion sequences of talking human heads from input audio signal. To capture the expressive, detailed nature of human heads, in…

3D Face AnimationAudio SynthesisMotion Synthesis

LinguaLinker: Audio-Driven Portraits Animation with Implicit Facial Control Enhancement

2024-07-26 · Rui Zhang, Yixiao Fang, Zhengnan Lu, Pei Cheng 외

This study delves into the intricacies of synchronizing facial dynamics with multilingual audio inputs, focusing on the creation of visually compelling, time-synchronized animations through diffusion-based techniques. Di…

Instruct-NeuralTalker: Editing Audio-Driven Talking Radiance Fields with Instructions

2023-06-19 · Yuqi Sun, Ruian He, Weimin Tan, Bo Yan

Recent neural talking radiance field methods have shown great success in photorealistic audio-driven talking face synthesis. In this paper, we propose a novel interactive framework that utilizes human instructions to edi…

Face GenerationTalking Face Generation

IMTalker: Efficient Audio-driven Talking Face Generation with Implicit Motion Transfer

2025-11-27 · Bo Chen, Tao Liu, Qi Chen, Xie Chen 외 arxiv

Talking face generation aims to synthesize realistic speaking portraits from a single image, yet existing methods often rely on explicit optical flow and local warping, which fail to model complex global motions and caus…

Talking Face Generation