paper-with-me

홈 › Papers

MotiMotion: Motion-Controlled Video Generation with Visual Reasoning

2026-05-21 · Lee Hsin-Ying, Hanwen Jiang, Yiqun Mei, Jing Shi, Ming-Hsuan Yang, Zhixin Shu arxiv

Current motion-controlled image-to-video generation models rigidly follow user-provided trajectories that are often sparse, imprecise, and causally incomplete. Such reliance often yields unnatural or implausible outcomes, especially by missing secondary causal consequences. To address this, we introduce MotiMotion, a novel framework that reformulates motion control as a reasoning-then-generation problem. To encourage causally grounded and commonsense-consistent interactions, we leverage a training-free vision-language reasoner to refine image-space coordinates of primary trajectories and to hallucinate plausible secondary motions. To further improve motion naturalness, we propose a confidence-aware control scheme that modulates guidance strength, enabling the model to closely follow high-confidence plans while correcting artifacts under low-confidence inputs with its internal generative priors. To support systematic evaluation, we curate a new image-to-video benchmark, MotiBench, consisting of interaction-centric scenes where new events are triggered by motion. Both VLM-based evaluation and a human study on MotiBench demonstrate that MotiMotion produces videos with more plausible object behaviors and interaction, and is preferred over existing approaches.

📄 PDF Abstract BibTeX arXiv:2605.22818

Code (0)

등록된 구현이 없습니다.

Tasks

Visual ReasoningVideo Generation

Similar Papers 제목 키워드 기반

MEAD: A Large-scale Audio-visual Dataset for Emotional Talking-face Generation

2020-08-01 · ECCV 2020 8 · Kaisiyuan Wang Qianyi Wu Linsen Song Zhuoqian Yang Wayne Wu Chen Qian Ran He Yu Qiao Chen Change Loy

The synthesis of natural emotional reactions is an essentialcriteria in vivid talking-face video generation. This criteria is nevertheless seldom taken into consideration in previous works due to the absence of a large-s…

Face GenerationTalking Face GenerationTalking Head GenerationVideo Generation

C-Drag: Chain-of-Thought Driven Motion Controller for Video Generation

2025-02-27 · Yuhao Li, Mirana Claire Angel, Salman Khan, Yu Zhu 외

Trajectory-based motion control has emerged as an intuitive and efficient approach for controllable video generation. However, the existing trajectory-based approaches are usually limited to only generating the motion tr…

ObjectVideo Generation

Syncphony: Synchronized Audio-to-Video Generation with Diffusion Transformers

2025-09-26 · Jibin Song, Mingi Kwon, Jaeseok Jeong, Youngjung Uh arxiv

Text-to-video and image-to-video generation have made rapid progress in visual quality, but they remain limited in controlling the precise timing of motion. In contrast, audio provides temporal cues aligned with video mo…

Video Generation

MagicMotion: Controllable Video Generation with Dense-to-Sparse Trajectory Guidance

2025-03-20 · Quanhao Li, Zhen Xing, Rui Wang, HUI ZHANG 외

Recent advances in video generation have led to remarkable improvements in visual quality and temporal coherence. Upon this, trajectory-controllable video generation has emerged to enable precise object motion control th…

Image to Video GenerationObjectVideo Generation

MotionClone: Training-Free Motion Cloning for Controllable Video Generation

2024-06-08 · Pengyang Ling, Jiazi Bu, Pan Zhang, Xiaoyi Dong 외

Motion-based controllable video generation offers the potential for creating captivating visual content. Existing methods typically necessitate model training to encode particular motion cues or incorporate fine-tuning t…

DenoisingMotion GenerationMotion SynthesisText-to-Video Generation+1