paper-with-me

홈 › Papers

IF-MDM: Implicit Face Motion Diffusion Model for High-Fidelity Realtime Talking Head Generation

2024-12-05 · Sejong Yang, Seoung Wug Oh, Yang Zhou, Seon Joo Kim

We introduce a novel approach for high-resolution talking head generation from a single image and audio input. Prior methods using explicit face models, like 3D morphable models (3DMM) and facial landmarks, often fall short in generating high-fidelity videos due to their lack of appearance-aware motion representation. While generative approaches such as video diffusion models achieve high video quality, their slow processing speeds limit practical application. Our proposed model, Implicit Face Motion Diffusion Model (IF-MDM), employs implicit motion to encode human faces into appearance-aware compressed facial latents, enhancing video generation. Although implicit motion lacks the spatial disentanglement of explicit models, which complicates alignment with subtle lip movements, we introduce motion statistics to help capture fine-grained motion information. Additionally, our model provides motion controllability to optimize the trade-off between motion intensity and visual quality during inference. IF-MDM supports real-time generation of 512x512 resolution videos at up to 45 frames per second (fps). Extensive evaluations demonstrate its superior performance over existing diffusion and explicit face models. The code will be released publicly, available alongside supplementary materials. The video results can be found on https://bit.ly/ifmdm_supplementary.

📄 PDF Abstract BibTeX arXiv:2412.04000

Code (0)

등록된 구현이 없습니다.

Tasks

DisentanglementTalking Head GenerationVideo Generation

Methods 이 논문이 사용한 방법론

Diffusion Diffusion models generate samples by gradually removing noise from a signal, and their training objective can be expressed as a reweighted variational lower-bound…

Similar Papers 제목 키워드 기반

Navigating Large-Pose Challenge for High-Fidelity Face Reenactment with Video Diffusion Model

2025-07-22 · Mingtao Guo, Guanyu Xing, Yanci Zhang, Yanli Liu arxiv

Face reenactment aims to generate realistic talking head videos by transferring motion from a driving video to a static source image while preserving the source identity. Although existing methods based on either implici…

Geometry-guided Emotion Modulation for Controllable and Photorealistic Emotional Talking Face Generation

2026-08-01 · Chenggong Hu, Shaoyin Ma, Yi Wang, Li Sun 외 arxiv

Audio-driven emotional talking face generation aims to synthesize realistic videos with expressive facial dynamics. However, existing methods struggle to balance controllability and visual fidelity. Although implicit rep…

Talking Face GenerationContinuous Control

High-fidelity and Lip-synced Talking Face Synthesis via Landmark-based Diffusion Model

2024-08-10 · Weizhi Zhong, Junfan Lin, Peixin Chen, Liang Lin 외

Audio-driven talking face video generation has attracted increasing attention due to its huge industrial potential. Some previous methods focus on learning a direct mapping from audio to visual content. Despite progress,…

Face GenerationTalking Face GenerationVideo Generation

Point-to-Point: Sparse Motion Guidance for Controllable Video Editing

2025-11-23 · Yeji Song, Jaehyun Lee, Mijin Koo, JunHoo Lee 외 arxiv

Accurately preserving motion while editing a subject remains a core challenge in video editing tasks. Existing methods often face a trade-off between edit and motion fidelity, as they rely on motion representations that …

Multimodal-driven Talking Face Generation via a Unified Diffusion-based Generator

2023-05-04 · Chao Xu, Shaoting Zhu, Junwei Zhu, Tianxin Huang 외

Multimodal-driven talking face generation refers to animating a portrait with the given pose, expression, and gaze transferred from the driving image and video, or estimated from the text and audio. However, existing met…

DenoisingFace GenerationFace SwappingTalking Face Generation