paper-with-me

Papers

Audio-driven Talking Face Video Generation with Learning-based Personalized Head Pose

2020-02-24 · Ran Yi, Zipeng Ye, Juyong Zhang, Hujun Bao, Yong-Jin Liu

Real-world talking faces often accompany with natural head movement. However, most existing talking face video generation methods only consider facial animation with fixed head pose. In this paper, we address this problem by proposing a deep neural network model that takes an audio signal A of a source person and a very short video V of a target person as input, and outputs a synthesized high-quality talking face video with personalized head pose (making use of the visual information in V), expression and lip synchronization (by considering both A and V). The most challenging issue in our work is that natural poses often cause in-plane and out-of-plane head rotations, which makes synthesized talking face video far from realistic. To address this challenge, we reconstruct 3D face animation and re-render it into synthesized frames. To fine tune these frames into realistic ones with smooth background transition, we propose a novel memory-augmented GAN module. By first training a general mapping based on a publicly available dataset and fine-tuning the mapping using the input short video of target person, we develop an effective strategy that only requires a small number of frames (about 300 frames) to learn personalized talking behavior including head pose. Extensive experiments and two user studies show that our method can generate high-quality (i.e., personalized head movements, expressions and good lip synchronization) talking face videos, which are naturally looking with more distinguishing head movement effects than the state-of-the-art methods.

📄 PDF Abstract BibTeX arXiv:2002.10137

Code (1)

yiranran/Audio-driven-TalkingFace-HeadPose 공식 구현 pytorch

Tasks

3D Face AnimationVideo Generation

Methods 이 논문이 사용한 방법론

Convolution A convolution is a type of matrix operation, consisting of a kernel, a small matrix of weights, that slides over input data performing element-wise multiplication with the…
Dogecoin Customer Service Number +1-833-534-1729 설명 없음

Similar Papers 제목 키워드 기반

JoyGen: Audio-Driven 3D Depth-Aware Talking-Face Video Editing

2025-01-03 · Qili Wang, Dajiang Wu, Zihang Xu, Junshi Huang 외

Significant progress has been made in talking-face video generation research; however, precise lip-audio synchronization and high visual quality remain challenging in editing lip shapes based on input audio. This paper i…

3D ReconstructionFace GenerationMotion GenerationTalking Face Generation+2

Audio-Driven Talking Face Video Generation with Dynamic Convolution Kernels

2022-01-16 · Zipeng Ye, Mengfei Xia, Ran Yi, Juyong Zhang 외

In this paper, we present a dynamic convolution kernel (DCK) strategy for convolutional neural networks. Using a fully convolutional network with the proposed DCKs, high-quality talking-face video can be generated from m…

Video Generation

Speech Driven Talking Face Generation from a Single Image and an Emotion Condition

2020-08-08 · Sefik Emre Eskimez, You Zhang, Zhiyao Duan

Visual emotion expression plays an important role in audiovisual speech communication. In this work, we propose a novel approach to rendering visual emotion expression in speech-driven talking face generation. Specifical…

Emotion RecognitionFace GenerationTalking Face Generation

One-shot Talking Face Generation from Single-speaker Audio-Visual Correlation Learning

2021-12-06 · Suzhen Wang, Lincheng Li, Yu Ding, Xin Yu

Audio-driven one-shot talking face generation methods are usually trained on video resources of various persons. However, their created videos often suffer unnatural mouth shapes and asynchronous lips because those metho…

Face GenerationTalking Face Generation

PortraitTalk: Towards Customizable One-Shot Audio-to-Talking Face Generation

2024-12-10 · Fatemeh Nazarieh, ZhenHua Feng, Diptesh Kanojia, Muhammad Awais 외

Audio-driven talking face generation is a challenging task in digital communication. Despite significant progress in the area, most existing methods concentrate on audio-lip synchronization, often overlooking aspects suc…

Face GenerationTalking Face Generation