paper-with-me

Papers

Thinking with Geometry: Active Geometry Integration for Spatial Reasoning

2026-02-05 · Haoyuan Li, Qihang Cao, Tao Tang, Kun Xiang, Zihan Guo, Jianhua Han, JiaWang Bian, Hang Xu, Xiaodan Liang arxiv

Recent progress in spatial reasoning with Multimodal Large Language Models (MLLMs) increasingly leverages geometric priors from 3D encoders. However, most existing integration strategies remain passive: geometry is exposed as a global stream and fused in an indiscriminate manner, which often induces semantic-geometry misalignment and redundant signals. We propose GeoThinker, a framework that shifts the paradigm from passive fusion to active perception. Instead of feature mixing, GeoThinker enables the model to selectively retrieve geometric evidence conditioned on its internal reasoning demands. GeoThinker achieves this through Spatial-Grounded Fusion applied at carefully selected VLM layers, where semantic visual priors selectively query and integrate task-relevant geometry via frame-strict cross-attention, further calibrated by Importance Gating that biases per-frame attention toward task-relevant structures. Comprehensive evaluation results show that GeoThinker sets a new state-of-the-art in spatial intelligence, achieving a peak score of 72.6 on the VSI-Bench. Furthermore, GeoThinker demonstrates robust generalization and significantly improved spatial perception across complex downstream scenarios, including embodied referring and autonomous driving. Our results indicate that the ability to actively integrate spatial structures is essential for next-generation spatial intelligence. Code can be found at https://github.com/Li-Hao-yuan/GeoThinker.

📄 PDF Abstract BibTeX arXiv:2602.06037

Code (0)

등록된 구현이 없습니다.

Tasks

Autonomous DrivingSpatial Reasoning

Similar Papers 제목 키워드 기반

ActiveNeRF: Learning Accurate 3D Geometry by Active Pattern Projection

2024-08-13 · Jianyu Tao, Changping Hu, Edward Yang, Jing Xu 외

NeRFs have achieved incredible success in novel view synthesis. However, the accuracy of the implicit geometry is unsatisfactory because the passive static environmental illumination has low spatial frequency and cannot …

3D geometryNeRFNovel View Synthesis

3DThinkVLA: Endowing Vision-Language-Action Models with Latent 3D Priors via 3D-Thinking-Guided Co-training

2026-06-03 · Jiaxin Shi, Xidong Zhang, Fucai Zhu, Zhe Li 외 arxiv

We propose a 3D-thinking-guided co-training framework that enables vision-language-action (VLA) models to perform 3D spatial reasoning implicitly during action prediction. Our core insight is that 3D geometry perception …

Spatial ReasoningText Generation

Influence of the Geometry of the world model on Curiosity Based Exploration

2023-04-01 · Grégoire Sergeant-Perthuis, Nils Ruet, David Rudrauf, Dimitri Ognibene 외

In human spatial awareness, 3-D projective geometry structures information integration and action planning through perspective taking within an internal representation space. The way different perspectives are related an…

SpatialStack: Layered Geometry-Language Fusion for 3D VLM Spatial Reasoning

2026-03-28 · Jian Zhang, Shijie Zhou, Bangya Liu, Achuta Kadambi 외 arxiv

Large vision-language models (VLMs) still struggle with reliable 3D spatial reasoning, a core capability for embodied and physical AI systems. This limitation arises from their inability to capture fine-grained 3D geomet…

Spatial Reasoning

Towards Scalable Multi-View Reconstruction of Geometry and Materials

2023-06-06 · Carolin Schmitt, Božidar Antić, Andrei Neculai, Joo Ho Lee 외

In this paper, we propose a novel method for joint recovery of camera pose, object geometry and spatially-varying Bidirectional Reflectance Distribution Function (svBRDF) of 3D scenes that exceed object-scale and hence c…

Distributed Optimization