paper-with-me

홈 › Papers

Anchored Diffusion for Video Face Reenactment

2024-07-21 · Idan Kligvasser, Regev Cohen, George Leifman, Ehud Rivlin, Michael Elad

Video generation has drawn significant interest recently, pushing the development of large-scale models capable of producing realistic videos with coherent motion. Due to memory constraints, these models typically generate short video segments that are then combined into long videos. The merging process poses a significant challenge, as it requires ensuring smooth transitions and overall consistency. In this paper, we introduce Anchored Diffusion, a novel method for synthesizing relatively long and seamless videos. We extend Diffusion Transformers (DiTs) to incorporate temporal information, creating our sequence-DiT (sDiT) model for generating short video segments. Unlike previous works, we train our model on video sequences with random non-uniform temporal spacing and incorporate temporal information via external guidance, increasing flexibility and allowing it to capture both short and long-term relationships. Furthermore, during inference, we leverage the transformer architecture to modify the diffusion process, generating a batch of non-uniform sequences anchored to a common frame, ensuring consistency regardless of temporal distance. To demonstrate our method, we focus on face reenactment, the task of creating a video from a source image that replicates the facial expressions and movements from a driving video. Through comprehensive experiments, we show our approach outperforms current techniques in producing longer consistent high-quality videos while offering editing capabilities.

📄 PDF Abstract BibTeX arXiv:2407.15153

Code (0)

등록된 구현이 없습니다.

Tasks

Face ReenactmentVideo Generation

Methods 이 논문이 사용한 방법론

Focus 설명 없음
Diffusion Diffusion models generate samples by gradually removing noise from a signal, and their training objective can be expressed as a reweighted variational lower-bound…

Similar Papers 제목 키워드 기반

TongueReenact: Geometry-Anchored Tongue Synthesis for Face Reenactment

2026-07-30 · MD Wahiduzzaman Khan, Mingshan Jia, Xiaolin Zhang, En Yu 외 arxiv

Modern face reenactment systems achieve impressive pose and expression transfer using geometry-driven representations. However, they largely ignore tongue dynamics, leading to anatomically inconsistent mouth interiors du…

DiffusionAct: Controllable Diffusion Autoencoder for One-shot Face Reenactment

2024-03-25 · Stella Bounareli, Christos Tzelepis, Vasileios Argyriou, Ioannis Patras 외

Video-driven neural face reenactment aims to synthesize realistic facial images that successfully preserve the identity and appearance of a source face, while transferring the target head pose and facial expressions. Exi…

Face ReenactmentImage Generation

Navigating Large-Pose Challenge for High-Fidelity Face Reenactment with Video Diffusion Model

2025-07-22 · Mingtao Guo, Guanyu Xing, Yanci Zhang, Yanli Liu arxiv

Face reenactment aims to generate realistic talking head videos by transferring motion from a driving video to a static source image while preserving the source identity. Although existing methods based on either implici…

Automatic Face Reenactment

2016-02-08 · CVPR 2014 6 · Pablo Garrido, Levi Valgaerts, Ole Rehmsen, Thorsten Thormaehlen 외

We propose an image-based, facial reenactment system that replaces the face of an actor in an existing target video with the face of a user from a source video, while preserving the original target performance. Our syste…

ClusteringFace ModelFace ReenactmentFace Transfer+2

TALK-Act: Enhance Textural-Awareness for 2D Speaking Avatar Reenactment with Diffusion Model

2024-10-14 · Jiazhi Guan, Quanwei Yang, Kaisiyuan Wang, Hang Zhou 외

Recently, 2D speaking avatars have increasingly participated in everyday scenarios due to the fast development of facial animation techniques. However, most existing works neglect the explicit control of human bodies. In…