paper-with-me

Papers

Facial Keypoint Sequence Generation from Audio

2020-11-02 · Prateek Manocha, Prithwijit Guha

Whenever we speak, our voice is accompanied by facial movements and expressions. Several recent works have shown the synthesis of highly photo-realistic videos of talking faces, but they either require a source video to drive the target face or only generate videos with a fixed head pose. This lack of facial movement is because most of these works focus on the lip movement in sync with the audio while assuming the remaining facial keypoints' fixed nature. To address this, a unique audio-keypoint dataset of over 150,000 videos at 224p and 25fps is introduced that relates the facial keypoint movement for the given audio. This dataset is then further used to train the model, Audio2Keypoint, a novel approach for synthesizing facial keypoint movement to go with the audio. Given a single image of the target person and an audio sequence (in any language), Audio2Keypoint generates a plausible keypoint movement sequence in sync with the input audio, conditioned on the input image to preserve the target person's facial characteristics. To the best of our knowledge, this is the first work that proposes an audio-keypoint dataset and learns a model to output the plausible keypoint sequence to go with audio of any arbitrary length. Audio2Keypoint generalizes across unseen people with a different facial structure allowing us to generate the sequence with the voice from any source or even synthetic voices. Instead of learning a direct mapping from audio to video domain, this work aims to learn the audio-keypoint mapping that allows for in-plane and out-of-plane head rotations, while preserving the person's identity using a Pose Invariant (PIV) Encoder.

📄 PDF Abstract BibTeX arXiv:2011.01114

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

FONT: Flow-guided One-shot Talking Head Generation with Natural Head Motions

2023-03-31 · Jin Liu, Xi Wang, Xiaomeng Fu, Yesheng Chai 외

One-shot talking head generation has received growing attention in recent years, with various creative and practical applications. An ideal natural and vivid generated talking head video should contain natural head pose …

DiversityPose PredictionTalking Head GenerationUnsupervised Keypoints

Audio2Head: Audio-driven One-shot Talking-head Generation with Natural Head Motion

2021-07-20 · Suzhen Wang, Lincheng Li, Yu Ding, Changjie Fan 외

We propose an audio-driven talking-head method to generate photo-realistic talking-head videos from a single reference image. In this work, we tackle two key challenges: (i) producing natural head motions that match spee…

Image GenerationTalking Head Generation

Unlock Pose Diversity: Accurate and Efficient Implicit Keypoint-based Spatiotemporal Diffusion for Audio-driven Talking Portrait

2025-03-17 · Chaolong Yang, Kai Yao, Yuyao Yan, Chenru Jiang 외

Audio-driven single-image talking portrait generation plays a crucial role in virtual reality, digital human creation, and filmmaking. Existing approaches are generally categorized into keypoint-based and image-based met…

Computational EfficiencyDiversity

Multi Modal Adaptive Normalization for Audio to Video Generation

2020-12-14 · Neeraj Kumar, Srishti Goel, Ankur Narang, Brejesh lall

Speech-driven facial video generation has been a complex problem due to its multi-modal aspects namely audio and video domain. The audio comprises lots of underlying features such as expression, pitch, loudness, prosody(…

Optical Flow EstimationSSIMVideo Generation

Audio-Driven Talking Face Generation with Diverse yet Realistic Facial Animations

2023-04-18 · Rongliang Wu, Yingchen Yu, Fangneng Zhan, Jiahui Zhang 외

Audio-driven talking face generation, which aims to synthesize talking faces with realistic facial animations (including accurate lip movements, vivid facial expression details and natural head poses) corresponding to th…

Face GenerationTalking Face Generation