paper-with-me

홈 › Papers

SpatialPIN: Enhancing Spatial Reasoning Capabilities of Vision-Language Models through Prompting and Interacting 3D Priors

2024-03-18 · Chenyang Ma, Kai Lu, Ta-Ying Cheng, Niki Trigoni, Andrew Markham

Current state-of-the-art spatial reasoning-enhanced VLMs are trained to excel at spatial visual question answering (VQA). However, we believe that higher-level 3D-aware tasks, such as articulating dynamic scene changes and motion planning, require a fundamental and explicit 3D understanding beyond current spatial VQA datasets. In this work, we present SpatialPIN, a framework designed to enhance the spatial reasoning capabilities of VLMs through prompting and interacting with priors from multiple 3D foundation models in a zero-shot, training-free manner. Extensive experiments demonstrate that our spatial reasoning-imbued VLM performs well on various forms of spatial VQA and can extend to help in various downstream robotics tasks such as pick and stack and trajectory planning.

📄 PDF Abstract BibTeX arXiv:2403.13438

Code (0)

등록된 구현이 없습니다.

Tasks

HallucinationMotion PlanningQuestion AnsweringSpatial ReasoningTrajectory PlanningVisual Question AnsweringVisual Question Answering (VQA)

Similar Papers 제목 키워드 기반

Sparkle: Mastering Basic Spatial Capabilities in Vision Language Models Elicits Generalization to Composite Spatial Reasoning

2024-10-21 · Yihong Tang, Ao Qu, Zhaokai Wang, Dingyi Zhuang 외

Vision language models (VLMs) have demonstrated impressive performance across a wide range of downstream tasks. However, their proficiency in spatial reasoning remains limited, despite its crucial role in tasks involving…

Spatial ReasoningSynthetic Data Generation

SpatiaLQA: A Benchmark for Evaluating Spatial Logical Reasoning in Vision-Language Models

2026-02-24 · Yuechen Xie, Xiaoyan Zhang, Yicheng Shan, Hao Zhu 외 arxiv

Vision-Language Models (VLMs) have been increasingly applied in real-world scenarios due to their outstanding understanding and reasoning capabilities. Although VLMs have already demonstrated impressive capabilities in c…

Visual Question AnsweringLogical Reasoning

Beyond Semantics: Rediscovering Spatial Awareness in Vision-Language Models

2025-03-21 · Jianing Qi, Jiawei Liu, Hao Tang, Zhigang Zhu

Vision-Language Models (VLMs) excel at identifying and describing objects but struggle with spatial reasoning such as accurately understanding the relative positions of objects. Inspired by the dual-pathway (ventral-dors…

DiagnosticObject RecognitionSpatial Reasoning

DepthVLA: Enhancing Vision-Language-Action Models with Depth-Aware Spatial Reasoning

2025-10-15 · Tianyuan Yuan, Yicheng Liu, Chenhao Lu, Zhuoguang Chen 외 arxiv

Vision-Language-Action (VLA) models have recently shown impressive generalization and language-guided manipulation capabilities. However, their performance degrades on tasks requiring precise spatial reasoning due to lim…

Spatial Reasoning

Enhancing Spatial Reasoning in Vision-Language Models via Chain-of-Thought Prompting and Reinforcement Learning

2025-07-06 · Binbin Ji, Siddharth Agrawal, Qiance Tang, Yvonne Wu arxiv

This study investigates the spatial reasoning capabilities of vision-language models (VLMs) through Chain-of-Thought (CoT) prompting and reinforcement learning. We begin by evaluating the impact of different prompting st…

Reinforcement LearningSpatial Reasoning