paper-with-me

Papers

StyleTalker: One-shot Style-based Audio-driven Talking Head Video Generation

2022-08-23 · Dongchan Min, Minyoung Song, Eunji Ko, Sung Ju Hwang

We propose StyleTalker, a novel audio-driven talking head generation model that can synthesize a video of a talking person from a single reference image with accurately audio-synced lip shapes, realistic head poses, and eye blinks. Specifically, by leveraging a pretrained image generator and an image encoder, we estimate the latent codes of the talking head video that faithfully reflects the given audio. This is made possible with several newly devised components: 1) A contrastive lip-sync discriminator for accurate lip synchronization, 2) A conditional sequential variational autoencoder that learns the latent motion space disentangled from the lip movements, such that we can independently manipulate the motions and lip movements while preserving the identity. 3) An auto-regressive prior augmented with normalizing flow to learn a complex audio-to-motion multi-modal latent space. Equipped with these components, StyleTalker can generate talking head videos not only in a motion-controllable way when another motion source video is given but also in a completely audio-driven manner by inferring realistic motions from the input audio. Through extensive experiments and user studies, we show that our model is able to synthesize talking head videos with impressive perceptual quality which are accurately lip-synced with the input audios, largely outperforming state-of-the-art baselines.

📄 PDF Abstract BibTeX arXiv:2208.10922

Code (0)

등록된 구현이 없습니다.

Tasks

Talking Head GenerationVideo Generation

Similar Papers 제목 키워드 기반

One-shot Talking Face Generation from Single-speaker Audio-Visual Correlation Learning

2021-12-06 · Suzhen Wang, Lincheng Li, Yu Ding, Xin Yu

Audio-driven one-shot talking face generation methods are usually trained on video resources of various persons. However, their created videos often suffer unnatural mouth shapes and asynchronous lips because those metho…

Face GenerationTalking Face Generation

OmniTalker: Real-Time Text-Driven Talking Head Generation with In-Context Audio-Visual Style Replication

2025-04-03 · Zhongjian Wang, Peng Zhang, Jinwei Qi, Guangyuan Wang Sheng Xu 외

Recent years have witnessed remarkable advances in talking head generation, owing to its potential to revolutionize the human-AI interaction from text interfaces into realistic video chats. However, research on text-driv…

Talking Head GenerationVideo Synchronization

DiffTED: One-shot Audio-driven TED Talk Video Generation with Diffusion-based Co-speech Gestures

2024-09-11 · Steven Hogue, Chenxu Zhang, Hamza Daruger, Yapeng Tian 외

Audio-driven talking video generation has advanced significantly, but existing methods often depend on video-to-video translation techniques and traditional generative networks like GANs and they typically generate takin…

DiversityTalking Head GenerationVideo Generation

Imitating Arbitrary Talking Style for Realistic Audio-DrivenTalking Face Synthesis

2021-10-30 · Haozhe Wu, Jia Jia, Haoyu Wang, Yishun Dou 외

People talk with diversified styles. For one piece of speech, different talking styles exhibit significant differences in the facial and head pose movements. For example, the "excited" style usually talks with the mouth …

Face Generation

PortraitTalk: Towards Customizable One-Shot Audio-to-Talking Face Generation

2024-12-10 · Fatemeh Nazarieh, ZhenHua Feng, Diptesh Kanojia, Muhammad Awais 외

Audio-driven talking face generation is a challenging task in digital communication. Despite significant progress in the area, most existing methods concentrate on audio-lip synchronization, often overlooking aspects suc…

Face GenerationTalking Face Generation