paper-with-me

Papers

Talking Face Generation by Adversarially Disentangled Audio-Visual Representation

2018-07-20 · Hang Zhou, Yu Liu, Ziwei Liu, Ping Luo, Xiaogang Wang

Talking face generation aims to synthesize a sequence of face images that correspond to a clip of speech. This is a challenging task because face appearance variation and semantics of speech are coupled together in the subtle movements of the talking face regions. Existing works either construct specific face appearance model on specific subjects or model the transformation between lip motion and speech. In this work, we integrate both aspects and enable arbitrary-subject talking face generation by learning disentangled audio-visual representation. We find that the talking face sequence is actually a composition of both subject-related information and speech-related information. These two spaces are then explicitly disentangled through a novel associative-and-adversarial training process. This disentangled representation has an advantage where both audio and video can serve as inputs for generation. Extensive experiments show that the proposed approach generates realistic talking face sequences on arbitrary subjects with much clearer lip motion patterns than previous work. We also demonstrate the learned audio-visual representation is extremely useful for the tasks of automatic lip reading and audio-video retrieval.

📄 PDF Abstract BibTeX arXiv:1807.07860

Code (1)

Hangz-nju-cuhk/Talking-Face-Generation-DAVS pytorch

Tasks

Face GenerationLip ReadingRetrievalTalking Face GenerationVideo Retrieval

Similar Papers 제목 키워드 기반

Animating Face using Disentangled Audio Representations

2019-10-02 · Gaurav Mittal, Baoyuan Wang

All previous methods for audio-driven talking head generation assume the input audio to be clean with a neutral tone. As we show empirically, one can easily break these systems by simply adding certain background noise t…

Representation LearningTalking Head Generation

JEAN: Joint Expression and Audio-guided NeRF-based Talking Face Generation

2024-09-18 · Sai Tanmay Reddy Chakkera, Aggelina Chatziagapi, Dimitris Samaras

We introduce a novel method for joint expression and audio-guided talking face generation. Recent approaches either struggle to preserve the speaker identity or fail to produce faithful facial expressions. To address the…

Contrastive LearningFace GenerationNeRFTalking Face Generation

MoDiTalker: Motion-Disentangled Diffusion Model for High-Fidelity Talking Head Generation

2024-03-28 · Seyeon Kim, Siyoon Jin, JiHye Park, Kihong Kim 외

Conventional GAN-based models for talking head generation often suffer from limited quality and unstable training. Recent approaches based on diffusion models aimed to address these limitations and improve fidelity. Howe…

Talking Head Generation

DFA-NeRF: Personalized Talking Head Generation via Disentangled Face Attributes Neural Rendering

2022-01-03 · Shunyu Yao, RuiZhe Zhong, Yichao Yan, Guangtao Zhai 외

While recent advances in deep neural networks have made it possible to render high-quality images, generating photo-realistic and personalized talking head remains challenging. With given audio, the key to tackling this …

NeRFNeural RenderingTalking Head Generation

FaceChain-ImagineID: Freely Crafting High-Fidelity Diverse Talking Faces from Disentangled Audio

2024-03-04 · CVPR 2024 1 · Chao Xu, Yang Liu, Jiazheng Xing, Weida Wang 외

In this paper, we abstract the process of people hearing speech, extracting meaningful cues, and creating various dynamically audio-consistent talking faces, termed Listening and Imagining, into the task of high-fidelity…

Disentanglement