paper-with-me

홈 › Papers

Theory of Space: Can Foundation Models Construct Spatial Beliefs through Active Exploration?

2026-02-04 · Pingyue Zhang, Zihan Huang, Yue Wang, Jieyu Zhang, Letian Xue, Zihan Wang, Qineng Wang, Keshigeyan Chandrasegaran, Ruohan Zhang, Yejin Choi, Ranjay Krishna, Jiajun Wu, Li Fei-Fei, Manling Li arxiv

Spatial embodied intelligence requires agents to act to acquire information under partial observability. While multimodal foundation models excel at passive perception, their capacity for active, self-directed exploration remains understudied. We propose Theory of Space, defined as an agent's ability to actively acquire information through self-directed, active exploration and to construct, revise, and exploit a spatial belief from sequential, partial observations. We evaluate this through a benchmark where the goal is curiosity-driven exploration to build an accurate cognitive map. A key innovation is spatial belief probing, which prompts models to reveal their internal spatial representations at each step. Our evaluation of state-of-the-art models reveals several critical bottlenecks. First, we identify an Active-Passive Gap, where performance drops significantly when agents must autonomously gather information. Second, we find high inefficiency, as models explore unsystematically compared to program-based proxies. Through belief probing, we diagnose that while perception is an initial bottleneck, global beliefs suffer from instability that causes spatial knowledge to degrade over time. Finally, using a false belief paradigm, we uncover Belief Inertia, where agents fail to update obsolete priors with new evidence. This issue is present in text-based agents but is particularly severe in vision-based models. Our findings suggest that current foundation models struggle to maintain coherent, revisable spatial beliefs during active exploration.

📄 PDF Abstract BibTeX arXiv:2602.07055

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

TimeToM: Temporal Space is the Key to Unlocking the Door of Large Language Models' Theory-of-Mind

2024-07-01 · Guiyang Hou, Wenqi Zhang, Yongliang Shen, Linjuan Wu 외

Theory of Mind (ToM)-the cognitive ability to reason about mental states of ourselves and others, is the foundation of social interaction. Although ToM comes naturally to humans, it poses a significant challenge to even …

A semantic embedding space based on large language models for modelling human beliefs

2024-08-13 · Byunghwee Lee, Rachith Aiyappa, Yong-Yeol Ahn, Haewoon Kwak 외

Beliefs form the foundation of human cognition and decision-making, guiding our actions and social connections. A model encapsulating beliefs and their interrelationships is crucial for understanding their influence on o…

Decision MakingLanguage ModelingLanguage ModellingLarge Language Model

Metaprobability and Dempster-Shafer in Evidential Reasoning

2013-03-27 · Robert Fung, Chee Yee Chong

Evidential reasoning in expert systems has often used ad-hoc uncertainty calculi. Although it is generally accepted that probability theory provides a firm theoretical foundation, researchers have found some problems wit…

Memory-Augmented Theory of Mind Network

2023-01-17 · Dung Nguyen, Phuoc Nguyen, Hung Le, Kien Do 외

Social reasoning necessitates the capacity of theory of mind (ToM), the ability to contextualise and attribute mental states to others without having access to their internal cognitive structure. Recent machine learning …

Attribute

UserHarness: Harnessing User Minds for Stronger Agent Theory-of-Mind

2026-05-26 · Cheng Qian, Jiayu Liu, Heng Ji arxiv

Understanding what a user believes and intends is central to building effective agent assistants. This ability is often evaluated through Theory-of-Mind (ToM) tasks, where success requires reasoning from the user's persp…