paper-with-me

Papers

MoDiTalker: Motion-Disentangled Diffusion Model for High-Fidelity Talking Head Generation

2024-03-28 · Seyeon Kim, Siyoon Jin, JiHye Park, Kihong Kim, Jiyoung Kim, Jisu Nam, Seungryong Kim

Conventional GAN-based models for talking head generation often suffer from limited quality and unstable training. Recent approaches based on diffusion models aimed to address these limitations and improve fidelity. However, they still face challenges, including extensive sampling times and difficulties in maintaining temporal consistency due to the high stochasticity of diffusion models. To overcome these challenges, we propose a novel motion-disentangled diffusion model for high-quality talking head generation, dubbed MoDiTalker. We introduce the two modules: audio-to-motion (AToM), designed to generate a synchronized lip motion from audio, and motion-to-video (MToV), designed to produce high-quality head video following the generated motion. AToM excels in capturing subtle lip movements by leveraging an audio attention mechanism. In addition, MToV enhances temporal consistency by leveraging an efficient tri-plane representation. Our experiments conducted on standard benchmarks demonstrate that our model achieves superior performance compared to existing models. We also provide comprehensive ablation studies and user study results.

📄 PDF Abstract BibTeX arXiv:2403.19144

Code (1)

KU-CVLAB/MoDiTalker 공식 구현 pytorch

Tasks

Talking Head Generation

Methods 이 논문이 사용한 방법론

Diffusion Diffusion models generate samples by gradually removing noise from a signal, and their training objective can be expressed as a reweighted variational lower-bound…

Similar Papers 제목 키워드 기반

DEMO: Disentangled Motion Latent Flow Matching for Fine-Grained Controllable Talking Portrait Synthesis

2025-10-12 · Peiyin Chen, Zhuowei Yang, Hui Feng, Sheng Jiang 외 arxiv

Audio-driven talking-head generation has advanced rapidly with diffusion-based generative models, yet producing temporally coherent videos with fine-grained motion control remains challenging. We propose DEMO, a flow-mat…

DeX-Portrait: Disentangled and Expressive Portrait Animation via Explicit and Latent Motion Representations

2025-12-17 · Yuxiang Shi, Zhe Li, Yanwen Wang, Hao Zhu 외 arxiv

Portrait animation from a single source image and a driving video is a long-standing problem. Recent approaches tend to adopt diffusion-based image/video generation models for realistic and expressive animation. However,…

Video Generation

DNF: Unconditional 4D Generation with Dictionary-based Neural Fields

2024-12-06 · CVPR 2025 1 · Xinyi Zhang, Naiqi Li, Angela Dai

While remarkable success has been achieved through diffusion-based 3D generative models for shapes, 4D generative modeling remains challenging due to the complexity of object deformations over time. We propose DNF, a new…

Dictionary Learning

Motion Consistency Model: Accelerating Video Diffusion with Disentangled Motion-Appearance Distillation

2024-06-11 · Yuanhao Zhai, Kevin Lin, Zhengyuan Yang, Linjie Li 외

Image diffusion distillation achieves high-fidelity generation with very few sampling steps. However, applying these techniques directly to video diffusion often results in unsatisfactory frame quality due to the limited…

Disentangled Hierarchical VAE for 3D Human-Human Interaction Generation

2026-02-24 · Zichen Geng, Zeeshan Hayder, Bo Miao, Jian Liu 외 arxiv

Generating realistic 3D Human-Human Interaction (HHI) requires coherent modeling of the physical plausibility of the agents and their interaction semantics. Existing methods compress all motion information into a single …

Computational EfficiencyContrastive Learning