paper-with-me

홈 › Papers

Towards Spatial Trace with Reasoning in Vision-Language Models for Robotics

2025-12-15 · Enshen Zhou, Yibo Li, Jingkun An, Jiayuan Zhang, Shanyu Rong, Mengzhen Liu, Yi Han, Yuheng Ji, Huajie Tan, Jiawei He, Pengwei Wang, Zhongyuan Wang, Cheng Chi, Lu Sheng, Shanghang Zhang arxiv

Spatial tracing, as a fundamental embodied interaction ability for robots, is inherently challenging as it requires multi-step metric-grounded reasoning compounded with complex spatial referring and real-world metric measurement. However, existing methods struggle with this compositional task. To this end, we propose RoboTracer, a 3D-aware VLM that first achieves both 3D spatial referring and measuring via a universal spatial encoder and a regression-supervised decoder to enhance scale awareness during supervised fine-tuning (SFT). Moreover, RoboTracer advances multi-step metric-grounded reasoning via reinforcement fine-tuning (RFT) with metric-sensitive process rewards, supervising key intermediate perceptual cues to accurately generate spatial traces. To support SFT and RFT training, we introduce TraceSpatial, a large-scale dataset of 30M QA pairs, spanning outdoor/indoor/tabletop scenes and supporting complex reasoning processes (up to 9 steps). We further present TraceSpatial-Bench, a challenging benchmark filling the gap to evaluate spatial tracing. Experimental results show that RoboTracer surpasses baselines in spatial understanding, measuring, and referring, with an average success rate of 79.1%, and also achieves SOTA performance on TraceSpatial-Bench by a large margin, exceeding Gemini-2.5-Pro by 36% accuracy. Notably, RoboTracer can be integrated with various control policies to execute long-horizon, dynamic tasks across diverse robots (UR5, G1 humanoid) in cluttered real-world scenes. Please see the project page at https://zhoues.github.io/RoboTracer.

📄 PDF Abstract BibTeX arXiv:2512.13660

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

SpatialTraceGen: High-Fidelity Traces for Efficient VLM Spatial Reasoning Distillation

2025-10-28 · Gio Huh, Dhruv Sheth, Rayhan Zirvi, Frank Xiao arxiv

While Vision-Language Models (VLMs) excel in many areas, they struggle with complex spatial reasoning, which requires problem decomposition and strategic tool use. Fine-tuning smaller, more deployable models offers an ef…

Reinforcement LearningSpatial Reasoning

TraceVision: Trajectory-Aware Vision-Language Model for Human-Like Spatial Understanding

2026-02-23 · Fan Yang, Shurong Zheng, Hongyin Zhao, Yufei Zhan 외 arxiv

Recent Large Vision-Language Models (LVLMs) demonstrate remarkable capabilities in image understanding and natural language generation. However, current approaches focus predominantly on global image understanding, strug…

Trajectory PredictionScene UnderstandingLogical Reasoning

RoboSpatial: Teaching Spatial Understanding to 2D and 3D Vision-Language Models for Robotics

2024-11-25 · CVPR 2025 1 · Chan Hee Song, Valts Blukis, Jonathan Tremblay, Stephen Tyree 외

Spatial understanding is a crucial capability that enables robots to perceive their surroundings, reason about their environment, and interact with it meaningfully. In modern robotics, these capabilities are increasingly…

Robot ManipulationScene UnderstandingSpatial Reasoning

InternVLA-M1: A Spatially Guided Vision-Language-Action Framework for Generalist Robot Policy

2025-10-15 · Xinyi Chen, Yilun Chen, Yanwei Fu, Ning Gao 외 arxiv

We introduce InternVLA-M1, a unified framework for spatial grounding and robot control that advances instruction-following robots toward scalable, general-purpose intelligence. Its core idea is spatially guided vision-la…

Instruction FollowingSpatial Reasoning

SpatialVLM: Endowing Vision-Language Models with Spatial Reasoning Capabilities

2024-01-22 · CVPR 2024 1 · Boyuan Chen, Zhuo Xu, Sean Kirmani, Brian Ichter 외

Understanding and reasoning about spatial relationships is a fundamental capability for Visual Question Answering (VQA) and robotics. While Vision Language Models (VLM) have demonstrated remarkable performance in certain…

Question AnsweringSpatial ReasoningVisual Question AnsweringVisual Question Answering (VQA)