paper-with-me

홈 › Papers

Why Is Spatial Reasoning Hard for VLMs? An Attention Mechanism Perspective on Focus Areas

2025-03-03 · Shiqi Chen, Tongyao Zhu, Ruochen Zhou, Jinghan Zhang, Siyang Gao, Juan Carlos Niebles, Mor Geva, Junxian He, Jiajun Wu, Manling Li

Large Vision Language Models (VLMs) have long struggled with spatial reasoning tasks. Surprisingly, even simple spatial reasoning tasks, such as recognizing "under" or "behind" relationships between only two objects, pose significant challenges for current VLMs. In this work, we study the spatial reasoning challenge from the lens of mechanistic interpretability, diving into the model's internal states to examine the interactions between image and text tokens. By tracing attention distribution over the image through out intermediate layers, we observe that successful spatial reasoning correlates strongly with the model's ability to align its attention distribution with actual object locations, particularly differing between familiar and unfamiliar spatial relationships. Motivated by these findings, we propose ADAPTVIS based on inference-time confidence scores to sharpen the attention on highly relevant regions when confident, while smoothing and broadening the attention window to consider a wider context when confidence is lower. This training-free decoding method shows significant improvement (e.g., up to a 50 absolute point improvement) on spatial reasoning benchmarks such as WhatsUp and VSR with negligible cost. We make code and data publicly available for research purposes at https://github.com/shiqichen17/AdaptVis.

📄 PDF Abstract BibTeX arXiv:2503.01773

Code (1)

shiqichen17/adaptvis 공식 구현 pytorch

Tasks

Spatial Reasoning

Methods 이 논문이 사용한 방법론

Softmax The Softmax output function transforms a previous layer's output into a vector of probabilities. It is commonly used for multiclass classification. Given an input vector $x$…
Attention 설명 없음
ALIGN In the ALIGN method, visual and language representations are jointly trained from noisy image alt-text data. The image and text encoders are learned via contrastive loss…

Similar Papers 제목 키워드 기반

Grounding Isn't Knowing: Do VLMs Need Object Localization for Spatial Reasoning?

2026-08-24 · Xiwei Liu, Yulong Li, Xinlin Zhuang, Xuhui Li 외 arxiv

Vision-language models (VLMs) can answer spatial questions, yet the mechanisms connecting object grounding to spatial reasoning remain poorly understood. It is underexplored whether spatial reasoning internally requires …

Object LocalizationSpatial Reasoning

Imagine in Space: Exploring the Frontier of Spatial Intelligence and Reasoning Efficiency in Vision Language Models

2025-11-16 · Xiaoxing Lian, Aidong Yang, Jun Zhu, Peng Wang 외 arxiv

Large language models (LLMs) and vision language models (VLMs), such as DeepSeek R1,OpenAI o3, and Gemini 2.5 Pro, have demonstrated remarkable reasoning capabilities across logical inference, problem solving, and decisi…

Spatial ReasoningDecision Making

MetaSpatial: Reinforcing 3D Spatial Reasoning in VLMs for the Metaverse

2025-03-24 · Zhenyu Pan, Han Liu

We present MetaSpatial, the first reinforcement learning (RL)-based framework designed to enhance 3D spatial reasoning in vision-language models (VLMs), enabling real-time 3D scene generation without the need for hard-co…

Layout GenerationReinforcement Learning (RL)Scene GenerationSpatial Reasoning

I Know About "Up"! Enhancing Spatial Reasoning in Visual Language Models Through 3D Reconstruction

2024-07-19 · Zaiqiao Meng, Hao Zhou, Yifang Chen

Visual Language Models (VLMs) are essential for various tasks, particularly visual reasoning tasks, due to their robust multi-modal information integration, visual reasoning capabilities, and contextual awareness. Howeve…

3D ReconstructionSpatial ReasoningVisual Reasoning

Linear Mechanisms for Spatiotemporal Reasoning in Vision Language Models

2026-01-18 · Raphi Kang, Hongqiao Chen, Georgia Gkioxari, Pietro Perona arxiv

Spatio-temporal reasoning is a remarkable capability of Vision Language Models (VLMs), but the underlying mechanisms of such abilities remain largely opaque. We postulate that visual/geometrical and textual representatio…