paper-with-me

홈 › Papers

RoboSpatial: Teaching Spatial Understanding to 2D and 3D Vision-Language Models for Robotics

2024-11-25 · CVPR 2025 1 · Chan Hee Song, Valts Blukis, Jonathan Tremblay, Stephen Tyree, Yu Su, Stan Birchfield

Spatial understanding is a crucial capability that enables robots to perceive their surroundings, reason about their environment, and interact with it meaningfully. In modern robotics, these capabilities are increasingly provided by vision-language models. However, these models face significant challenges in spatial reasoning tasks, as their training data are based on general-purpose image datasets that often lack sophisticated spatial understanding. For example, datasets frequently do not capture reference frame comprehension, yet effective spatial reasoning requires understanding whether to reason from ego-, world-, or object-centric perspectives. To address this issue, we introduce RoboSpatial, a large-scale dataset for spatial understanding in robotics. It consists of real indoor and tabletop scenes, captured as 3D scans and egocentric images, and annotated with rich spatial information relevant to robotics. The dataset includes 1M images, 5k 3D scans, and 3M annotated spatial relationships, and the pairing of 2D egocentric images with 3D scans makes it both 2D- and 3D- ready. Our experiments show that models trained with RoboSpatial outperform baselines on downstream tasks such as spatial affordance prediction, spatial relationship prediction, and robot manipulation.

📄 PDF Abstract BibTeX arXiv:2411.16537

Code (0)

등록된 구현이 없습니다.

Tasks

Robot ManipulationScene UnderstandingSpatial Reasoning

Similar Papers 제목 키워드 기반

SpaceTools: Tool-Augmented Spatial Reasoning via Double Interactive RL

2025-12-03 · Siyi Chen, Mikaela Angelina Uy, Chan Hee Song, Faisal Ladhak 외 arxiv

Vision Language Models (VLMs) demonstrate strong qualitative visual understanding, but struggle with metrically precise spatial reasoning required for embodied applications. The agentic paradigm promises that VLMs can us…

Reinforcement LearningSpatial Reasoning

Technical Report of RoboSpatial Challenge at CVPR 2026: Selective Reasoning Activation and Reference-Frame Disambiguation for Embodied Spatial Reasoning

2026-06-30 · Yuxiang Xie, Qi Lv, Jianming Xing, Zijian Hong 외 arxiv

Vision-language models achieve strong general perception but often struggle with the spatial reasoning required for embodied tasks. We present RoboSpatialBrain, our submission to the RoboSpatial Challenge at the Embodied…

Spatial Reasoning

Learning Multi-View Spatial Reasoning from Cross-View Relations

2026-03-30 · Suchae Jeong, Jaehwi Song, Haeone Lee, Hanna Kim 외 arxiv

Vision-language models (VLMs) have achieved impressive results on single-view vision tasks, but lack the multi-view spatial reasoning capabilities essential for embodied AI systems to understand 3D environments and manip…

Spatial Reasoning

Loc3R-VLM: Language-based Localization and 3D Reasoning with Vision-Language Models

2026-03-18 · Kevin Qu, Haozhe Qi, Mihai Dusmanu, Mahdi Rad 외 arxiv

Multimodal Large Language Models (MLLMs) have made impressive progress in connecting vision and language, but they still struggle with spatial understanding and viewpoint-aware reasoning. Recent efforts aim to augment th…

From Flatland to Space: Teaching Vision-Language Models to Perceive and Reason in 3D

2025-03-29 · Jiahui Zhang, Yurui Chen, Yanpeng Zhou, Yueming Xu 외

Recent advances in LVLMs have improved vision-language understanding, but they still struggle with spatial perception, limiting their ability to reason about complex 3D scenes. Unlike previous approaches that incorporate…

Spatial Reasoning