paper-with-me

홈 › Papers

SpatialAct: Probing Spatial Reasoning-to-Action Capabilities of VLM Agents in 3D Scenes

2026-05-29 · Tianhui Liu, Jie Feng, Zhiheng Zheng, Shengyuan Wang, Yiming Guo, Yanxin Xi, Hangyu Fan, Yong Li, Pan Hui arxiv

Humans can effortlessly perceive spatial layouts, form cognitive representations, reason about spatial relations, and translate such reasoning into actions in everyday 3D environments. Although recent vision-language models (VLMs) have shown promising performance on observation-conditioned spatial perception and reasoning tasks, it remains unclear whether they can build coherent spatial understanding, act upon it, and refine their actions through multi-turn feedback. To study this problem, we introduce \textbf{SpatialAct}, a simulator-grounded benchmark for probing \textit{action-conditioned spatial reasoning} in 3D scenes. Starting from the most challenging setting, Multi-turn Interactive Refinement, we further design its decomposed counterpart, Single-step Error Detection and Fix, together with five fundamental spatial ability tasks to diagnose the underlying causes of model failures. Experiments reveal a clear reasoning-to-action gap: current VLMs can perform well on isolated spatial reasoning tasks, but struggle to maintain coherent spatial beliefs and produce reliable actions during multi-turn feedback, substantially underperforming humans. These results suggest that current VLM agents still lack robust spatial state tracking under action-induced environment changes, even when low-level control is abstracted away.

📄 PDF Abstract BibTeX arXiv:2605.31148

Code (0)

등록된 구현이 없습니다.

Tasks

Spatial Reasoning

Similar Papers 제목 키워드 기반

SpatialActor: Exploring Disentangled Spatial Representations for Robust Robotic Manipulation

2025-11-12 · Hao Shi, Bin Xie, Yingfei Liu, Yang Yue 외 arxiv

Robotic manipulation requires precise spatial understanding to interact with objects in the real world. Point-based methods suffer from sparse sampling, leading to the loss of fine-grained semantics. Image-based methods …

Perceive, Interact, Reason: Building Tool-Augmented Visual Agents for Spatial Reasoning

2026-06-11 · Changye Li, Meng Lu, Yi Wu, Ligeng Zhu arxiv

While recent vision-language models (VLMs) demonstrate strong multimodal understanding, they remain limited in spatial reasoning tasks that require active evidence acquisition and multi-step visual interaction. This limi…

Spatial Reasoning

Are Large Language Models Geospatially Knowledgeable?

2023-10-09 · Prabin Bhandari, Antonios Anastasopoulos, Dieter Pfoser

Despite the impressive performance of Large Language Models (LLM) for various natural language processing tasks, little is known about their comprehension of geographic data and related ability to facilitate informed geo…

Decision Making

Ask the World Before Acting: Environment Probing for Calibrated Agent World Models

2026-06-30 · Xinyuan Song, Zekun Cai arxiv

Language agents acting over long horizons must maintain beliefs about tool states, object locations, graph edges, and subgoal dependencies. When these beliefs drift, failures can be fixed neither by longer reasoning trac…

SpatialLLM: A Compound 3D-Informed Design towards Spatially-Intelligent Large Multimodal Models

2025-05-01 · CVPR 2025 1 · Wufei Ma, Luoxin Ye, Nessa McWeeney, Celso M de Melo 외

Humans naturally understand 3D spatial relationships, enabling complex reasoning like predicting collisions of vehicles from different directions. Current large multimodal models (LMMs), however, lack of this capability …

Spatial ReasoningVisual Question Answering (VQA)