paper-with-me

Papers

Geometrically-Constrained Agent for Spatial Reasoning

2025-11-27 · Zeren Chen, Xiaoya Lu, Zhijie Zheng, Pengrui Li, Lehan He, Yijin Zhou, Jing Shao, Bohan Zhuang, Lu Sheng arxiv

Vision Language Models (VLMs) exhibit a fundamental semantic-to-geometric gap in spatial reasoning: they excel at qualitative semantic inference but their reasoning operates within a lossy semantic space, misaligned with high-fidelity geometry. Current paradigms fail to bridge this gap. Training-based methods suffer from an ``oracle paradox,'' learning flawed spatial logic from imperfect oracles. Tool-integrated methods constrain the final computation but critically leave the VLM's planning process unconstrained, resulting in geometrically flawed plans. In this work, we propose Geometrically-Constrained Agent (GCA), a training-free agentic paradigm that resolves this gap by introducing a formal task constraint. Specifically, we strategically decouples the VLM's role into two stages. First, acting as a semantic analyst, the VLM translates the user's ambiguous query into the formal, verifiable task constraint, which defines the reference frame and objective. Second, acting as a task solver, the VLM generates and executes tool calls strictly within the deterministic bounds defined by the constraint. This geometrically-constrained reasoning strategy successfully resolve the semantic-to-geometric gap, yielding a robust and verifiable reasoning pathway for spatial reasoning. Comprehensive experiments demonstrate that GCA achieves SOTA performance on multiple spatial reasoning benchmarks, surpassing existing training-based and tool-integrated methods by ~27%. Please see our homepage at https://gca-spatial-reasoning.github.io.

📄 PDF Abstract BibTeX arXiv:2511.22659

Code (0)

등록된 구현이 없습니다.

Tasks

Spatial Reasoning

Similar Papers 제목 키워드 기반

Semantic Radiance Fields as Simulators for Spatial Reasoning in Real-World Scenes

2026-08-13 · Nico Heider, Michał Jan Włodarczyk, Katarzyna Wasielewska-Michniewska, Przemysław Hołda 외 arxiv

Training and evaluating spatial reasoning in embodied agents requires diverse environments that are both geometrically faithful and semantically queryable. Synthetic simulators offer ground truth semantics but sacrifice …

Spatial Reasoning

Boosting MLLM Spatial Reasoning with Geometrically Referenced 3D Scene Representations

2026-03-09 · Jiangye Yuan, Gowri Kumar, Baoyuan Wang arxiv

While Multimodal Large Language Models (MLLMs) have achieved remarkable success in 2D visual understanding, their ability to reason about 3D space remains limited. To address this gap, we introduce geometrically referenc…

Mathematical ReasoningSpatial Reasoning

Navigate Complex Physical Worlds via Geometrically Constrained LLM

2024-10-23 · Yongqiang Huang, Wentao Ye, Liyao Li, Junbo Zhao

This study investigates the potential of Large Language Models (LLMs) for reconstructing and constructing the physical world solely based on textual knowledge. It explores the impact of model performance on spatial under…

Navigate

OmniCoT: A Benchmark for Global and Multi-Step Panoramic Reasoning

2026-06-29 · Haocong He, Chenfei Liao, Zichen Wen, Zihao Dongfang 외 arxiv

Multimodal Large Language Models (MLLMs) have demonstrated promising spatial reasoning capabilities, while these abilities remain underexplored in the emerging visual modality of panoramic imagery. The full 360°$\times$1…

Spatial Reasoning

Holi-Spatial: Evolving Video Streams into Holistic 3D Spatial Intelligence

2026-03-08 · Yuanyuan Gao, Hao Li, Yifei Liu, Xinhao Ji 외 arxiv

The pursuit of spatial intelligence fundamentally relies on access to large-scale, fine-grained 3D data. However, existing approaches predominantly construct spatial understanding benchmarks by generating question-answer…

Spatial Reasoning