paper-with-me

Papers

SpatialVLM: Endowing Vision-Language Models with Spatial Reasoning Capabilities

2024-01-22 · CVPR 2024 1 · Boyuan Chen, Zhuo Xu, Sean Kirmani, Brian Ichter, Danny Driess, Pete Florence, Dorsa Sadigh, Leonidas Guibas, Fei Xia

Understanding and reasoning about spatial relationships is a fundamental capability for Visual Question Answering (VQA) and robotics. While Vision Language Models (VLM) have demonstrated remarkable performance in certain VQA benchmarks, they still lack capabilities in 3D spatial reasoning, such as recognizing quantitative relationships of physical objects like distances or size differences. We hypothesize that VLMs' limited spatial reasoning capability is due to the lack of 3D spatial knowledge in training data and aim to solve this problem by training VLMs with Internet-scale spatial reasoning data. To this end, we present a system to facilitate this approach. We first develop an automatic 3D spatial VQA data generation framework that scales up to 2 billion VQA examples on 10 million real-world images. We then investigate various factors in the training recipe, including data quality, training pipeline, and VLM architecture. Our work features the first internet-scale 3D spatial reasoning dataset in metric space. By training a VLM on such data, we significantly enhance its ability on both qualitative and quantitative spatial VQA. Finally, we demonstrate that this VLM unlocks novel downstream applications in chain-of-thought spatial reasoning and robotics due to its quantitative estimation capability. Project website: https://spatial-vlm.github.io/

📄 PDF Abstract BibTeX arXiv:2401.12168

Code (0)

등록된 구현이 없습니다.

Tasks

Question AnsweringSpatial ReasoningVisual Question AnsweringVisual Question Answering (VQA)

Similar Papers 제목 키워드 기반

PhysVLM: Enabling Visual Language Models to Understand Robotic Physical Reachability

2025-03-11 · CVPR 2025 1 · Weijie Zhou, Manli Tao, Chaoyang Zhao, Haiyun Guo 외

Understanding the environment and a robot's physical reachability is crucial for task execution. While state-of-the-art vision-language models (VLMs) excel in environmental perception, they often generate inaccurate or i…

Visual Reasoning

SpaAct: Spatially-Activated Transition Learning with Curriculum Adaptation for Vision-Language Navigation

2026-04-30 · Pengna Li, Kangyi Wu, Shaoqing Xu, Fang Li 외 arxiv

Vision-and-Language Navigation (VLN) aims to enable an embodied agent to follow natural-language instructions and navigate to a target location in unseen 3D environments. We argue that adapting VLMs to VLN requires endow…

Vision-Language Navigation

3DThinkVLA: Endowing Vision-Language-Action Models with Latent 3D Priors via 3D-Thinking-Guided Co-training

2026-06-03 · Jiaxin Shi, Xidong Zhang, Fucai Zhu, Zhe Li 외 arxiv

We propose a 3D-thinking-guided co-training framework that enables vision-language-action (VLA) models to perform 3D spatial reasoning implicitly during action prediction. Our core insight is that 3D geometry perception …

Spatial ReasoningText Generation

Endowing Embodied Agents with Spatial Reasoning Capabilities for Vision-and-Language Navigation

2025-04-09 · Luo Ling, Bai Qianqian

Enhancing the spatial perception capabilities of mobile robots is crucial for achieving embodied Vision-and-Language Navigation (VLN). Although significant progress has been made in simulated environments, directly trans…

HallucinationSpatial ReasoningVision and Language Navigation

3D-Generalist: Self-Improving Vision-Language-Action Models for Crafting 3D Worlds

2025-07-09 · Fan-Yun Sun, Shengguang Wu, Christian Jacobsen, Thomas Yim 외 arxiv

Despite large-scale pretraining endowing models with language and vision reasoning capabilities, improving their spatial reasoning capability remains challenging due to the lack of data grounded in the 3D world. While it…

Synthetic Data GenerationSpatial Reasoning