paper-with-me

Papers

MotionV2V: Editing Motion in a Video

2025-11-25 · Ryan Burgert, Charles Herrmann, Forrester Cole, Michael S Ryoo, Neal Wadhwa, Andrey Voynov, Nataniel Ruiz arxiv

While generative video models have achieved remarkable fidelity and consistency, applying these capabilities to video editing remains a complex challenge. Recent research has explored motion controllability as a means to enhance text-to-video generation or image animation; however, we identify precise motion control as a promising yet under-explored paradigm for editing existing videos. In this work, we propose modifying video motion by directly editing sparse trajectories extracted from the input. We term the deviation between input and output trajectories a "motion edit" and demonstrate that this representation, when coupled with a generative backbone, enables powerful video editing capabilities. To achieve this, we introduce a pipeline for generating "motion counterfactuals", video pairs that share identical content but distinct motion, and we fine-tune a motion-conditioned video diffusion architecture on this dataset. Our approach allows for edits that start at any timestamp and propagate naturally. In a four-way head-to-head user study, our model achieves over 65 percent preference against prior work. Please see our project page: https://ryanndagreat.github.io/MotionV2V

📄 PDF Abstract BibTeX arXiv:2511.20640

Code (0)

등록된 구현이 없습니다.

Tasks

Text-to-Video Generation

Similar Papers 제목 키워드 기반

MotionVerse: A Unified Multimodal Framework for Motion Comprehension, Generation and Editing

2025-09-28 · Ruibing Hou, Mingshuang Luo, Hongyu Pan, Hong Chang 외 arxiv

This paper proposes MotionVerse, a unified framework that harnesses the capabilities of Large Language Models (LLMs) to comprehend, generate, and edit human motion in both single-person and multi-person scenarios. To eff…

Computational Efficiency

MotionVLA: Injecting Geometric Motion into Vision-Language-Action Model

2026-06-06 · Shanglin Yuan, Weiheng Zhao, Xianda Guo, Wei Sui 외 arxiv

Vision-language-action (VLA) models increasingly condition robot policies on history, depth, or 4D features to resolve ambiguity in long-horizon manipulation. However, more spatiotemporal evidence is not necessarily bett…

MotionVideoGAN: A Novel Video Generator Based on the Motion Space Learned from Image Pairs

2023-03-06 · Jingyuan Zhu, Huimin Ma, Jiansheng Chen, Jian Yuan

Video generation has achieved rapid progress benefiting from high-quality renderings provided by powerful image generators. We regard the video synthesis task as generating a sequence of images sharing the same contents …

Motion GenerationUnconditional Video GenerationVideo Generation

MotionVLA: Vision-Language-Action Model for Humanoid Motion

2026-06-13 · Nonghai Zhang, Siyu Zhai, Yanjun Li, Zeyu Zhang 외 arxiv

Generating realistic humanoid motion from scene images and text involves both low-frequency pose semantics and high-frequency physical dynamics. However, many existing methods tokenize motion with a single shared codeboo…

Vanishing-Point-Guided Video Semantic Segmentation of Driving Scenes

2024-01-27 · CVPR 2024 1 · Diandian Guo, Deng-Ping Fan, Tongyu Lu, Christos Sakaridis 외

The estimation of implicit cross-frame correspondences and the high computational cost have long been major challenges in video semantic segmentation (VSS) for driving scenes. Prior works utilize keyframes, feature propa…

Motion EstimationSegmentationSemantic SegmentationVideo Semantic Segmentation