paper-with-me

홈 › Papers

Magic-Me: Identity-Specific Video Customized Diffusion

2024-02-14 · Ze Ma, Daquan Zhou, Chun-Hsiao Yeh, Xue-She Wang, Xiuyu Li, Huanrui Yang, Zhen Dong, Kurt Keutzer, Jiashi Feng

Creating content with specified identities (ID) has attracted significant interest in the field of generative models. In the field of text-to-image generation (T2I), subject-driven creation has achieved great progress with the identity controlled via reference images. However, its extension to video generation is not well explored. In this work, we propose a simple yet effective subject identity controllable video generation framework, termed Video Custom Diffusion (VCD). With a specified identity defined by a few images, VCD reinforces the identity characteristics and injects frame-wise correlation at the initialization stage for stable video outputs. To achieve this, we propose three novel components that are essential for high-quality identity preservation and stable video generation: 1) a noise initialization method with 3D Gaussian Noise Prior for better inter-frame stability; 2) an ID module based on extended Textual Inversion trained with the cropped identity to disentangle the ID information from the background 3) Face VCD and Tiled VCD modules to reinforce faces and upscale the video to higher resolution while preserving the identity's features. We conducted extensive experiments to verify that VCD is able to generate stable videos with better ID over the baselines. Besides, with the transferability of the encoded identity in the ID module, VCD is also working well with personalized text-to-image models available publicly. The codes are available at https://github.com/Zhen-Dong/Magic-Me.

📄 PDF Abstract BibTeX arXiv:2402.09368

Code (1)

zhen-dong/magic-me 공식 구현 pytorch

Tasks

Image GenerationText to Image GenerationText-to-Image GenerationVideo Generation

Methods 이 논문이 사용한 방법론

Diffusion Diffusion models generate samples by gradually removing noise from a signal, and their training objective can be expressed as a reweighted variational lower-bound…

Similar Papers 제목 키워드 기반

MagicID: Hybrid Preference Optimization for ID-Consistent and Dynamic-Preserved Video Customization

2025-03-16 · Hengjia Li, Lifan Jiang, Xi Xiao, Tianyang Wang 외

Video identity customization seeks to produce high-fidelity videos that maintain consistent identity and exhibit significant dynamics based on users' reference images. However, existing approaches face two key challenges…

Magic Mirror: ID-Preserved Video Generation in Video Diffusion Transformers

2025-01-07 · Yuechen Zhang, Yaoyang Liu, Bin Xia, Bohao Peng 외

We present Magic Mirror, a framework for generating identity-preserved videos with cinematic-level quality and dynamic motion. While recent advances in video diffusion models have shown impressive capabilities in text-to…

DiversityText-to-Video GenerationVideo Generation

MAGIC-Talk: Motion-aware Audio-Driven Talking Face Generation with Customizable Identity Control

2025-10-26 · Fatemeh Nazarieh, Zhenhua Feng, Diptesh Kanojia, Muhammad Awais 외 arxiv

Audio-driven talking face generation has gained significant attention for applications in digital media and virtual avatars. While recent methods improve audio-lip synchronization, they often struggle with temporal consi…

Talking Face GenerationVideo Generation

DreaMoving: A Human Video Generation Framework based on Diffusion Models

2023-12-08 · Mengyang Feng, Jinlin Liu, Kai Yu, Yuan YAO 외

In this paper, we present DreaMoving, a diffusion-based controllable video generation framework to produce high-quality customized human videos. Specifically, given target identity and posture sequences, DreaMoving can g…

Video Generation

MagicInfinite: Generating Infinite Talking Videos with Your Words and Voice

2025-03-07 · Hongwei Yi, Tian Ye, Shitong Shao, Xuancheng Yang 외

We present MagicInfinite, a novel diffusion Transformer (DiT) framework that overcomes traditional portrait animation limitations, delivering high-fidelity results across diverse character types-realistic humans, full-bo…

DenoisingPortrait AnimationVideo Generation