paper-with-me

Papers

MotionZero:Exploiting Motion Priors for Zero-shot Text-to-Video Generation

2023-11-28 · Sitong Su, Litao Guo, Lianli Gao, HengTao Shen, Jingkuan Song

Zero-shot Text-to-Video synthesis generates videos based on prompts without any videos. Without motion information from videos, motion priors implied in prompts are vital guidance. For example, the prompt "airplane landing on the runway" indicates motion priors that the "airplane" moves downwards while the "runway" stays static. Whereas the motion priors are not fully exploited in previous approaches, thus leading to two nontrivial issues: 1) the motion variation pattern remains unaltered and prompt-agnostic for disregarding motion priors; 2) the motion control of different objects is inaccurate and entangled without considering the independent motion priors of different objects. To tackle the two issues, we propose a prompt-adaptive and disentangled motion control strategy coined as MotionZero, which derives motion priors from prompts of different objects by Large-Language-Models and accordingly applies motion control of different objects to corresponding regions in disentanglement. Furthermore, to facilitate videos with varying degrees of motion amplitude, we propose a Motion-Aware Attention scheme which adjusts attention among frames by motion amplitude. Extensive experiments demonstrate that our strategy could correctly control motion of different objects and support versatile applications including zero-shot video edit.

📄 PDF Abstract BibTeX arXiv:2311.16635

Code (0)

등록된 구현이 없습니다.

Tasks

DisentanglementText-to-Video GenerationVideo GenerationZero-shot Text-to-Video Generation

Similar Papers 제목 키워드 기반

Motion-Zero: Zero-Shot Moving Object Control Framework for Diffusion-Based Video Generation

2024-01-18 · Changgu Chen, Junwei Shu, Gaoqi He, Changbo Wang 외

Recent large-scale pre-trained diffusion models have demonstrated a powerful generative ability to produce high-quality videos from detailed text descriptions. However, exerting control over the motion of objects in vide…

DenoisingPositionVideo Generation

IM-Zero: Instance-level Motion Controllable Video Generation in a Zero-shot Manner

2025-01-01 · CVPR 2025 1 · YuYang Huang, Yabo Chen, Li Ding, Xiaopeng Zhang 외

Controllability of video generation has been recently concerned in addition to the quality of generated videos. The main challenge to controllable video generation is to synthesize videos based on user-specified inst…

Motion GenerationText-to-Video GenerationVideo Generation

Imagine2Real: Towards Zero-shot Humanoid-Object Interaction via Video Generative Priors

2026-05-21 · Jiahe Chen, ZiRui Wang, Feiyu Jia, Xiao Chen 외 arxiv

Whole-body Humanoid-Object Interaction (HOI) is bottlenecked by the scarcity of high-fidelity 3D data. While video generative priors offer a promising alternative, existing methods suffer from \textit{Representation Misa…

PromptStereo: Zero-Shot Stereo Matching via Structure and Motion Prompts

2026-03-02 · Xianqi Wang, Hao Yang, Hangtian Wang, Junda Cheng 외 arxiv

Modern stereo matching methods have leveraged monocular depth foundation models to achieve superior zero-shot generalization performance. However, most existing methods primarily focus on extracting robust features for c…

Zero-shot Generalization

Generalizable Object Keypoint Localization from Generative Priors

2025-01-01 · CVPR 2025 1 · Dongkai Wang, Jiang Duan, Liangjian Wen, Shiyu Xuan 외

Generalizable object keypoint localization is a fundamental computer vision task in understanding the object structure. It is challenging for existing keypoint localization methods because their limited training data…

Cross-Domain Few-ShotImage GenerationObject