paper-with-me

Papers

DiT-Head: High-Resolution Talking Head Synthesis using Diffusion Transformers

2023-12-11 · Aaron Mir, Eduardo Alonso, Esther Mondragón

We propose a novel talking head synthesis pipeline called "DiT-Head", which is based on diffusion transformers and uses audio as a condition to drive the denoising process of a diffusion model. Our method is scalable and can generalise to multiple identities while producing high-quality results. We train and evaluate our proposed approach and compare it against existing methods of talking head synthesis. We show that our model can compete with these methods in terms of visual quality and lip-sync accuracy. Our results highlight the potential of our proposed approach to be used for a wide range of applications, including virtual assistants, entertainment, and education. For a video demonstration of the results and our user study, please refer to our supplementary material.

📄 PDF Abstract BibTeX arXiv:2312.06400

Code (0)

등록된 구현이 없습니다.

Tasks

Denoising

Methods 이 논문이 사용한 방법론

Diffusion Diffusion models generate samples by gradually removing noise from a signal, and their training objective can be expressed as a reweighted variational lower-bound…

Similar Papers 제목 키워드 기반

DiffTalk: Crafting Diffusion Models for Generalized Audio-Driven Portraits Animation

2023-01-10 · CVPR 2023 1 · Shuai Shen, Wenliang Zhao, Zibin Meng, Wanhua Li 외

Talking head synthesis is a promising approach for the video production industry. Recently, a lot of effort has been devoted in this research area to improve the generation quality or enhance the model generalization. Ho…

DenoisingTalking Head Generation

NeRF-3DTalker: Neural Radiance Field with 3D Prior Aided Audio Disentanglement for Talking Head Synthesis

2025-02-20 · Xiaoxing Liu, Zhilei Liu, Chongke Bi

Talking head synthesis is to synthesize a lip-synchronized talking head video using audio. Recently, the capability of NeRF to enhance the realism and texture details of synthesized talking heads has attracted the attent…

DisentanglementNeRF

SyncTalk: The Devil is in the Synchronization for Talking Head Synthesis

2023-11-29 · CVPR 2024 1 · Ziqiao Peng, Wentao Hu, Yue Shi, Xiangyu Zhu 외

Achieving high synchronization in the synthesis of realistic, speech-driven talking head videos presents a significant challenge. Traditional Generative Adversarial Networks (GAN) struggle to maintain consistent facial i…

NeRFTalking Face GenerationTalking Head Generation

DreamHead: Learning Spatial-Temporal Correspondence via Hierarchical Diffusion for Audio-driven Talking Head Synthesis

2024-09-16 · Fa-Ting Hong, Yunfei Liu, Yu Li, Changyin Zhou 외

Audio-driven talking head synthesis strives to generate lifelike video portraits from provided audio. The diffusion model, recognized for its superior quality and robust generalization, has been explored for this task. H…

Talking Head Generation

EmbedTalk: Triplane-Free Talking Head Synthesis using Embedding-Driven Gaussian Deformation

2026-03-08 · Arpita Saggar, Jonathan C. Darling, Duygu Sarikaya, David C. Hogg arxiv

Real-time talking head synthesis increasingly relies on deformable 3D Gaussian Splatting (3DGS) due to its low latency. Tri-planes are the standard choice for encoding Gaussians prior to deformation, since they provide a…