paper-with-me

Papers

Diffused Heads: Diffusion Models Beat GANs on Talking-Face Generation

2023-01-06 · Michał Stypułkowski, Konstantinos Vougioukas, Sen He, Maciej Zięba, Stavros Petridis, Maja Pantic

Talking face generation has historically struggled to produce head movements and natural facial expressions without guidance from additional reference videos. Recent developments in diffusion-based generative models allow for more realistic and stable data synthesis and their performance on image and video generation has surpassed that of other generative models. In this work, we present an autoregressive diffusion model that requires only one identity image and audio sequence to generate a video of a realistic talking human head. Our solution is capable of hallucinating head movements, facial expressions, such as blinks, and preserving a given background. We evaluate our model on two different datasets, achieving state-of-the-art results on both of them.

📄 PDF Abstract BibTeX arXiv:2301.03396

Code (0)

등록된 구현이 없습니다.

Tasks

Face GenerationTalking Face GenerationVideo Generation

Methods 이 논문이 사용한 방법론

Diffusion Diffusion models generate samples by gradually removing noise from a signal, and their training objective can be expressed as a reweighted variational lower-bound…

Similar Papers 제목 키워드 기반

Diffusion-GAN: Training GANs with Diffusion

2022-06-05 · Zhendong Wang, Huangjie Zheng, Pengcheng He, Weizhu Chen 외

Generative adversarial networks (GANs) are challenging to train stably, and a promising remedy of injecting instance noise into the discriminator input has not been very effective in practice. In this paper, we propose D…

Image Generation

DiffTED: One-shot Audio-driven TED Talk Video Generation with Diffusion-based Co-speech Gestures

2024-09-11 · Steven Hogue, Chenxu Zhang, Hamza Daruger, Yapeng Tian 외

Audio-driven talking video generation has advanced significantly, but existing methods often depend on video-to-video translation techniques and traditional generative networks like GANs and they typically generate takin…

DiversityTalking Head GenerationVideo Generation

ScanTalk: 3D Talking Heads from Unregistered Scans

2024-03-16 · Federico Nocentini, Thomas Besnier, Claudio Ferrari, Sylvain Arguillere 외

Speech-driven 3D talking heads generation has emerged as a significant area of interest among researchers, presenting numerous challenges. Existing methods are constrained by animating faces with fixed topologies, wherei…

AI killed the video star. Audio-driven diffusion model for expressive talking head generation

2025-11-27 · Baptiste Chopin, Tashvik Dhamija, Pranav Balaji, Yaohui Wang 외 arxiv

We propose Dimitra++, a novel framework for audio-driven talking head generation, streamlined to learn lip motion, facial expression, as well as head pose motion. Specifically, we propose a conditional Motion Diffusion T…

Talking Head Generation

GaussianHeadTalk: Wobble-Free 3D Talking Heads with Audio Driven Gaussian Splatting

2025-12-11 · Madhav Agarwal, Mingtian Zhang, Laura Sevilla-Lara, Steven McDonagh arxiv

Speech-driven talking heads have recently emerged and enable interactive avatars. However, real-world applications are limited, as current methods achieve high visual fidelity but slow or fast yet temporally unstable. Di…

Image Generation