paper-with-me

Papers

MotionVLA: Injecting Geometric Motion into Vision-Language-Action Model

2026-06-06 · Shanglin Yuan, Weiheng Zhao, Xianda Guo, Wei Sui, Li Yu, Wenyu Liu, Xinggang Wang arxiv

Vision-language-action (VLA) models increasingly condition robot policies on history, depth, or 4D features to resolve ambiguity in long-horizon manipulation. However, more spatiotemporal evidence is not necessarily better: when the injected evidence is not motion-consistent, it can introduce geometric drift, fragmented temporal cues, and unstable action generation. This raises a simple question: should a VLA remember past frames, or remember the motion that connects them? We introduce MotionVLA, a motion-history interface that converts a short past-only video window into compact, time-continuous trajectory-field tokens. Instead of treating history as a sparse set of ndependently lifted frames, MotionVLA represents recent observations as physically coherent motion evidence. Current visual tokens query this history to retrieve task-relevant motion information, which is then recoupled into the VLA stream under trajectory-grounded supervision. Experiments across simulation benchmarks and preliminary real-robot rollouts show that MotionVLA improves long-horizon manipulation while producing smoother and more direct executions. These results suggest that effective VLA memory is not just about providing more 4D context, but about exposing motion-consistent evidence that is usable for control.

📄 PDF Abstract BibTeX arXiv:2606.08288

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

MotionVLA: Vision-Language-Action Model for Humanoid Motion

2026-06-13 · Nonghai Zhang, Siyu Zhai, Yanjun Li, Zeyu Zhang 외 arxiv

Generating realistic humanoid motion from scene images and text involves both low-frequency pose semantics and high-frequency physical dynamics. However, many existing methods tokenize motion with a single shared codeboo…

UniGeo: Unifying Geometric Guidance for Camera-Controllable Image Editing via Video Models

2026-04-19 · Hong Jiang, Wensong Song, Zongxin Yang, Ruijie Quan 외 arxiv

Camera-controllable image editing aims to synthesize novel views of a given scene under varying camera poses while strictly preserving cross-view geometric consistency. However, existing methods typically rely on fragmen…

Image EditingPoint Clouds

DRoPS: Dynamic 3D Reconstruction of Pre-Scanned Objects

2026-03-25 · Narek Tumanyan, Samuel Rota Bulò, Denis Rozumny, Lorenzo Porzi 외 arxiv

Dynamic scene reconstruction from casual videos has seen recent remarkable progress. Numerous approaches have attempted to overcome the ill-posedness of the task by distilling priors from 2D foundational models and by im…

3D Reconstruction

TGRIP: A Text-Guided Approach to Vehicle Instance Prediction in Autonomous Driving

2026-07-06 · Miguel Antunes-García, Santiago Montiel-Marín, Fabio Sánchez-García, Rodrigo Gutiérrez-Moreno 외 arxiv

Bird's-Eye View (BEV) end-to-end instance prediction has emerged as a robust paradigm for autonomous driving perception, effectively mitigating the error propagation inherent in traditional modular pipelines. However, cu…

Autonomous Driving

Beyond 3D VQAs: Injecting 3D Spatial Priors into Vision-Language Models for Enhanced Geometric Reasoning

2026-05-28 · Chun-Hsiao Yeh, Shengyi Qian, Manchen Wang, Yi Ma 외 arxiv

Vision-Language Models (VLMs) often struggle with robust 3D spatial reasoning. Prevailing methods that rely on fine-tuning with 3D visual question-answering (VQA) datasets may overfit dataset-specific biases, while integ…

Spatial Reasoning