paper-with-me

홈 › Papers

World2Mind: Cognition Toolkit for Allocentric Spatial Reasoning in Foundation Models

2026-03-10 · Shouwei Ruan, Bin Wang, Zhenyu Wu, Qihui Zhu, Yuxiang Zhang, Hang Su, Yubin Wang arxiv

Achieving robust spatial reasoning remains a fundamental challenge for current Multimodal Foundation Models (MFMs). Existing methods either overfit statistical shortcuts via 3D grounding data or remain confined to 2D visual perception, limiting both spatial reasoning accuracy and generalization in unseen scenarios. Inspired by the spatial cognitive mapping mechanisms of biological intelligence, we propose World2Mind, a training-free spatial intelligence toolkit. At its core, World2Mind leverages 3D reconstruction and instance segmentation models to construct structured spatial cognitive maps, empowering MFMs to proactively acquire targeted spatial knowledge regarding interested landmarks and routes of interest. To provide robust geometric-topological priors, World2Mind synthesizes an Allocentric-Spatial Tree (AST) that uses elliptical parameters to model the top-down layout of landmarks accurately. To mitigate the inherent inaccuracies of 3D reconstruction, we introduce a three-stage reasoning chain comprising tool invocation assessment, modality-decoupled cue collection, and geometry-semantics interwoven reasoning. Extensive experiments demonstrate that World2Mind boosts the performance of frontier models, such as GPT-5.2, by 5%~18%. Astonishingly, relying solely on the AST-structured text, purely text-only foundation models can perform complex 3D spatial reasoning, achieving performance approaching that of advanced multimodal models.

📄 PDF Abstract BibTeX arXiv:2603.09774

Code (0)

등록된 구현이 없습니다.

Tasks

Instance SegmentationSpatial Reasoning3D Reconstruction

Similar Papers 제목 키워드 기반

AlloSpatial: Agentic Harness Framework for Spatial Reasoning in Foundation Models

2026-06-08 · Shouwei Ruan, Bin Wang, Zhenyu Wu, Qihui Zhu 외 arxiv

Multimodal Foundation Models (MFMs) have made substantial progress, yet remain fragile in spatial reasoning over the physical world. A key bottleneck lies in their inability to transform local egocentric observations int…

Reinforcement LearningSpatial Reasoning

SpaceMind++: Toward Allocentric Cognitive Maps for Spatially Grounded Video MLLMs

2026-05-10 · Bo Gu, Zhikang Zhang, Zizhuang Wei, Zhenyuan Chen 외 arxiv

Recent multimodal large language models (MLLMs) have made remarkable progress in visual understanding and language-based reasoning, yet they lack a persistent world-centered representation for spatially consistent reason…

Allocentric Perceiver: Disentangling Allocentric Reasoning from Egocentric Visual Priors via Frame Instantiation

2026-02-05 · Hengyi Wang, Ruiqiang Zhang, Chang Liu, Guanjie Wang 외 arxiv

With the rising need for spatially grounded tasks such as Vision-Language Navigation/Action, allocentric perception capabilities in Vision-Language Models (VLMs) are receiving growing focus. However, VLMs remain brittle …

Vision-Language NavigationSpatial Reasoning

Keep it SymPL: Symbolic Projective Layout for Allocentric Spatial Reasoning in Vision-Language Models

2026-02-22 · Jaeyun Jang, Seunghui Shin, Taeho Park, Hyoseok Hwang arxiv

Perspective-aware spatial reasoning involves understanding spatial relationships from specific viewpoints-either egocentric (observer-centered) or allocentric (object-centered). While vision-language models (VLMs) perfor…

Spatial Reasoning

Conversational Orientation Reasoning: Egocentric-to-Allocentric Navigation with Multimodal Chain-of-Thought

2025-09-20 · Yu Ti Huang arxiv

Conversational agents must translate egocentric utterances (e.g., "on my right") into allocentric orientations (N/E/S/W). This challenge is particularly critical in indoor or complex facilities where GPS signals are weak…

Spatial Reasoning