MoDiT: Learning Highly Consistent 3D Motion Coefficients with Diffusion Transformer for Talking Head Generation
Audio-driven talking head generation is critical for applications such as virtual assistants, video games, and films, where natural lip movements are essential. Despite progress in this field, challenges remain in producing both consistent and realistic facial animations. Existing methods, often based on GANs or UNet-based diffusion models, face three major limitations: (i) temporal jittering caused by weak temporal constraints, resulting in frame inconsistencies; (ii) identity drift due to insufficient 3D information extraction, leading to poor preservation of facial identity; and (iii) unnatural blinking behavior due to inadequate modeling of realistic blink dynamics. To address these issues, we propose MoDiT, a novel framework that combines the 3D Morphable Model (3DMM) with a Diffusion-based Transformer. Our contributions include: (i) A hierarchical denoising strategy with revised temporal attention and biased self/cross-attention mechanisms, enabling the model to refine lip synchronization and progressively enhance full-face coherence, effectively mitigating temporal jittering. (ii) The integration of 3DMM coefficients to provide explicit spatial constraints, ensuring accurate 3D-informed optical flow prediction and improved lip synchronization using Wav2Lip results, thereby preserving identity consistency. (iii) A refined blinking strategy to model natural eye movements, with smoother and more realistic blinking behaviors.
Code (0)
등록된 구현이 없습니다.
Tasks
Talking Head GenerationInformation ExtractionSimilar Papers 제목 키워드 기반
MoDiTalker: Motion-Disentangled Diffusion Model for High-Fidelity Talking Head Generation
Conventional GAN-based models for talking head generation often suffer from limited quality and unstable training. Recent approaches based on diffusion models aimed to address these limitations and improve fidelity. Howe…
Talking Head GenerationLocal Analysis of Heterogeneous Intracellular Transport: Slow and Fast moving Endosomes
Trajectories of endosomes inside living eukaryotic cells are highly heterogeneous in space and time and diffuse anomalously due to a combination of viscoelasticity, caging, aggregation and active transport. Some of the t…
FD2Talk: Towards Generalized Talking Head Generation with Facial Decoupled Diffusion Model
Talking head generation is a significant research topic that still faces numerous challenges. Previous works often adopt generative adversarial networks or regression models, which are plagued by generation quality and a…
Talking Head GenerationConsistent and Controllable Image Animation with Motion Diffusion Models
Diffusion models have achieved significant progress in the task of image animation due to their powerful generative capabilities. However, preserving appearance consistency to the static input image, and avoiding abr…
Image AnimationVideo EditingUnravelling Heterogeneous Transport of Endosomes
A major open problem in biophysics is to understand the highly heterogeneous transport of many structures inside living cells, such as endosomes. We find that mathematically it is described by spatio-temporal heterogeneo…