paper-with-me

Papers

VideoComposer: Compositional Video Synthesis with Motion Controllability

2023-06-03 · NeurIPS 2023 11 · Xiang Wang, Hangjie Yuan, Shiwei Zhang, Dayou Chen, Jiuniu Wang, Yingya Zhang, Yujun Shen, Deli Zhao, Jingren Zhou

The pursuit of controllability as a higher standard of visual content creation has yielded remarkable progress in customizable image synthesis. However, achieving controllable video synthesis remains challenging due to the large variation of temporal dynamics and the requirement of cross-frame temporal consistency. Based on the paradigm of compositional generation, this work presents VideoComposer that allows users to flexibly compose a video with textual conditions, spatial conditions, and more importantly temporal conditions. Specifically, considering the characteristic of video data, we introduce the motion vector from compressed videos as an explicit control signal to provide guidance regarding temporal dynamics. In addition, we develop a Spatio-Temporal Condition encoder (STC-encoder) that serves as a unified interface to effectively incorporate the spatial and temporal relations of sequential inputs, with which the model could make better use of temporal conditions and hence achieve higher inter-frame consistency. Extensive experimental results suggest that VideoComposer is able to control the spatial and temporal patterns simultaneously within a synthesized video in various forms, such as text description, sketch sequence, reference video, or even simply hand-crafted motions. The code and models will be publicly available at https://videocomposer.github.io.

📄 PDF Abstract BibTeX arXiv:2306.02018

Code (4)

ali-vilab/i2vgen-xl pytorch
ali-vilab/videocomposer pytorch
damo-vilab/videocomposer pytorch
mindspore-lab/mindone mindspore

Tasks

Image GenerationText-to-Video Generation

Similar Papers 제목 키워드 기반

ReImagine: Rethinking Controllable High-Quality Human Video Generation via Image-First Synthesis

2026-04-21 · Zhengwentai Sun, Keru Zheng, Chenghong Li, Hongjie Liao 외 arxiv

Human video generation remains challenging due to the difficulty of jointly modeling human appearance, motion, and camera viewpoint under limited multi-view data. Existing methods often address these factors separately, …

Video GenerationImage Generation

Directing the World: Fast Autoregressive Video Generation with Compositional Human-Camera Control

2026-06-26 · Haoyuan Wang, Yabo Chen, Haibin Huang, Chi Zhang 외 arxiv

Building interactive world models requires generating realistic videos while maintaining controllable dynamics over long horizons. Autoregressive video generation offers a scalable foundation, but suffers from error accu…

Video Generation

BlobGEN-Vid: Compositional Text-to-Video Generation with Blob Video Representations

2025-01-13 · CVPR 2025 1 · Weixi Feng, Chao Liu, Sifei Liu, William Yang Wang 외

Existing video generation models struggle to follow complex text prompts and synthesize multiple objects, raising the need for additional grounding input for improved controllability. In this work, we propose to decompos…

ObjectText-to-Video GenerationVideo Generation

I2VControl: Disentangled and Unified Video Motion Synthesis Control

2024-11-26 · Wanquan Feng, Tianhao Qi, Jiawei Liu, Mingzhen Sun 외

Video synthesis techniques are undergoing rapid progress, with controllability being a significant aspect of practical usability for end-users. Although text condition is an effective way to guide video synthesis, captur…

Motion Synthesis

HECTOR: Hybrid Editable Compositional Object References for Video Generation

2026-03-09 · Guofeng Zhang, Angtian Wang, Jacob Zhiyuan Fang, Liming Jiang 외 arxiv

Real-world videos naturally portray complex interactions among distinct physical objects, effectively forming dynamic compositions of visual elements. However, most current video generation models synthesize scenes holis…

Video Generation