paper-with-me

Papers

Audio2Head: Audio-driven One-shot Talking-head Generation with Natural Head Motion

2021-07-20 · Suzhen Wang, Lincheng Li, Yu Ding, Changjie Fan, Xin Yu

We propose an audio-driven talking-head method to generate photo-realistic talking-head videos from a single reference image. In this work, we tackle two key challenges: (i) producing natural head motions that match speech prosody, and (ii) maintaining the appearance of a speaker in a large head motion while stabilizing the non-face regions. We first design a head pose predictor by modeling rigid 6D head movements with a motion-aware recurrent neural network (RNN). In this way, the predicted head poses act as the low-frequency holistic movements of a talking head, thus allowing our latter network to focus on detailed facial movement generation. To depict the entire image motions arising from audio, we exploit a keypoint based dense motion field representation. Then, we develop a motion field generator to produce the dense motion fields from input audio, head poses, and a reference image. As this keypoint based representation models the motions of facial regions, head, and backgrounds integrally, our method can better constrain the spatial and temporal consistency of the generated videos. Finally, an image generation network is employed to render photo-realistic talking-head videos from the estimated keypoint based motion fields and the input reference image. Extensive experiments demonstrate that our method produces videos with plausible head motions, synchronized facial expressions, and stable backgrounds and outperforms the state-of-the-art.

📄 PDF Abstract BibTeX arXiv:2107.09293

Code (1)

wangsuzhen/Audio2Head 공식 구현 pytorch

Tasks

Image GenerationTalking Head Generation

Similar Papers 제목 키워드 기반

StyleTalker: One-shot Style-based Audio-driven Talking Head Video Generation

2022-08-23 · Dongchan Min, Minyoung Song, Eunji Ko, Sung Ju Hwang

We propose StyleTalker, a novel audio-driven talking head generation model that can synthesize a video of a talking person from a single reference image with accurately audio-synced lip shapes, realistic head poses, and …

Talking Head GenerationVideo Generation

NeRFFaceSpeech: One-shot Audio-driven 3D Talking Head Synthesis via Generative Prior

2024-05-09 · Gihoon Kim, Kwanggyoon Seo, Sihun Cha, Junyong Noh

Audio-driven talking head generation is advancing from 2D to 3D content. Notably, Neural Radiance Field (NeRF) is in the spotlight as a means to synthesize high-quality 3D talking head outputs. Unfortunately, this NeRF-b…

Face ModelNeRFTalking Head Generation

AE-NeRF: Audio Enhanced Neural Radiance Field for Few Shot Talking Head Synthesis

2023-12-18 · Dongze Li, Kang Zhao, Wei Wang, Bo Peng 외

Audio-driven talking head synthesis is a promising topic with wide applications in digital human, film making and virtual reality. Recent NeRF-based approaches have shown superiority in quality and fidelity compared to p…

Face GenerationNeRFTalking Head Generation

S^3D-NeRF: Single-Shot Speech-Driven Neural Radiance Field for High Fidelity Talking Head Synthesis

2024-08-18 · Dongze Li, Kang Zhao, Wei Wang, Yifeng Ma 외

Talking head synthesis is a practical technique with wide applications. Current Neural Radiance Field (NeRF) based approaches have shown their superiority on driving one-shot talking heads with videos or signals regresse…

NeRF

Text-to-Video: a Two-stage Framework for Zero-shot Identity-agnostic Talking-head Generation

2023-08-12 · Zhichao Wang, Mengyu Dai, Keld Lundgaard

The advent of ChatGPT has introduced innovative methods for information gathering and analysis. However, the information provided by ChatGPT is limited to text, and the visualization of this information remains constrain…

Talking Head Generationtext-to-speechText to Speech