paper-with-me

Papers

Motion-I2V: Consistent and Controllable Image-to-Video Generation with Explicit Motion Modeling

2024-01-29 · Xiaoyu Shi, Zhaoyang Huang, Fu-Yun Wang, Weikang Bian, Dasong Li, Yi Zhang, Manyuan Zhang, Ka Chun Cheung, Simon See, Hongwei Qin, Jifeng Dai, Hongsheng Li

We introduce Motion-I2V, a novel framework for consistent and controllable image-to-video generation (I2V). In contrast to previous methods that directly learn the complicated image-to-video mapping, Motion-I2V factorizes I2V into two stages with explicit motion modeling. For the first stage, we propose a diffusion-based motion field predictor, which focuses on deducing the trajectories of the reference image's pixels. For the second stage, we propose motion-augmented temporal attention to enhance the limited 1-D temporal attention in video latent diffusion models. This module can effectively propagate reference image's feature to synthesized frames with the guidance of predicted trajectories from the first stage. Compared with existing methods, Motion-I2V can generate more consistent videos even at the presence of large motion and viewpoint variation. By training a sparse trajectory ControlNet for the first stage, Motion-I2V can support users to precisely control motion trajectories and motion regions with sparse trajectory and region annotations. This offers more controllability of the I2V process than solely relying on textual instructions. Additionally, Motion-I2V's second stage naturally supports zero-shot video-to-video translation. Both qualitative and quantitative comparisons demonstrate the advantages of Motion-I2V over prior approaches in consistent and controllable image-to-video generation. Please see our project page at https://xiaoyushi97.github.io/Motion-I2V/.

📄 PDF Abstract BibTeX arXiv:2401.15977

Code (0)

등록된 구현이 없습니다.

Tasks

Image to Video GenerationVideo Generation

Methods 이 논문이 사용한 방법론

Diffusion Diffusion models generate samples by gradually removing noise from a signal, and their training objective can be expressed as a reweighted variational lower-bound…

Similar Papers 제목 키워드 기반

Control-A-Video: Controllable Text-to-Video Diffusion Models with Motion Prior and Reward Feedback Learning

2023-05-23 · Weifeng Chen, Yatai Ji, Jie Wu, Hefeng Wu 외

Recent advances in text-to-image (T2I) diffusion models have enabled impressive image generation capabilities guided by text prompts. However, extending these techniques to video generation remains challenging, with exis…

Image GenerationOptical Flow EstimationStyle TransferText-to-Video Generation+3

IM-Zero: Instance-level Motion Controllable Video Generation in a Zero-shot Manner

2025-01-01 · CVPR 2025 1 · YuYang Huang, Yabo Chen, Li Ding, Xiaopeng Zhang 외

Controllability of video generation has been recently concerned in addition to the quality of generated videos. The main challenge to controllable video generation is to synthesize videos based on user-specified inst…

Motion GenerationText-to-Video GenerationVideo Generation

Track the Noise, Move the World:3D-Grounded Motion-Consistent Noise for Controllable Video Generation

2026-07-02 · Long Vu, Tan Ngo, Animesh Karnewar, Amir Habibian 외 arxiv

Modern image-and-text-to-video diffusion models can synthesize highly realistic videos by iteratively denoising an initial Gaussian noise tensor conditioned on reference image and text inputs. However, existing approache…

Video Generation

MOFA-Video: Controllable Image Animation via Generative Motion Field Adaptions in Frozen Image-to-Video Diffusion Model

2024-05-30 · Muyao Niu, Xiaodong Cun, Xintao Wang, Yong Zhang 외

We present MOFA-Video, an advanced controllable image animation method that generates video from the given image using various additional controllable signals (such as human landmarks reference, manual trajectories, and …

Image AnimationVideo Generation

ReImagine: Rethinking Controllable High-Quality Human Video Generation via Image-First Synthesis

2026-04-21 · Zhengwentai Sun, Keru Zheng, Chenghong Li, Hongjie Liao 외 arxiv

Human video generation remains challenging due to the difficulty of jointly modeling human appearance, motion, and camera viewpoint under limited multi-view data. Existing methods often address these factors separately, …

Video GenerationImage Generation