paper-with-me

홈 › Papers

ESI-Bench: Towards Embodied Spatial Intelligence that Closes the Perception-Action Loop

2026-05-18 · Yining Hong, Jiageng Liu, Han Yin, Manling Li, Leonidas Guibas, Li Fei-Fei, Jiajun Wu, Yejin Choi arxiv

Spatial intelligence unfolds through a perception-action loop: agents act to acquire observations, and reason about how observations vary as a function of action. Rather than passively processing what is seen, they actively uncover what is unseen - occluded structure, dynamics, containment, and functionality that cannot be resolved from passive sensing alone. We move beyond prior formulations of spatial intelligence that assume oracle observations by recasting the observer as an actor. We introduce ESI-BENCH, a comprehensive benchmark for embodied spatial intelligence spanning 10 task categories and 29 subcategories built on OmniGibson, grounded in Spelke's core knowledge systems. Agents must decide what abilities to deploy - perception, locomotion, and manipulation - and how to sequence them to actively accumulate task-relevant evidence. We conduct extensive experiments on state-of-the-art MLLMs and find that active exploration substantially outperforms passive counterparts, with agents spontaneously discovering emergent spatial strategies without explicit instructions, while random multi-view often adds noise rather than signal despite consuming far more images. Most failures stem not from weak perception but from action blindness: poor action choices lead to poor observations, which in turn drive cascading errors. While explicit 3D grounding stabilizes reasoning on depth-sensitive tasks, imperfect 3D representation proves more harmful than 2D baselines by distorting spatial relations. Human studies further reveal that unlike humans who seek falsifying viewpoints and revise beliefs under contradiction, models commit prematurely with high confidence regardless of evidence quality, exposing a metacognitive gap that neither better perception nor more embodied interaction alone can close.

📄 PDF Abstract BibTeX arXiv:2605.18746

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Embodied3DBench: Benchmarking Low-Level Embodied Spatial Intelligence of Vision Language Models

2026-05-27 · Jiyao Zhang, Mingxu Zhang, Yitong Peng, Haoxuan Liu 외 arxiv

Are current Vision Language Models (VLMs) ready to comprehend and reason about complex embodied interactions in 3D environments? We introduce Embodied3DBench, a robot-centric benchmark targeting low-level spatial intelli…

Trajectory PredictionSpatial Reasoning

Extending Embodied Question Answering from Perception to Decision

2026-05-25 · Xicheng Gong, Qiwei Li, Peiran Xu, Yadong Mu arxiv

Embodied Question Answering (EQA) connects perception, reasoning, and interaction within embodied environments. However, existing datasets and benchmarks remain fragmented, each focusing on a limited subset of reasoning …

Question Answering

NavSpace: How Navigation Agents Follow Spatial Intelligence Instructions

2025-10-09 · Haolin Yang, Yuxing Long, Zhuoyuan Yu, Zihan Yang 외 arxiv

Instruction-following navigation is a key step toward embodied intelligence. Prior benchmarks mainly focus on semantic understanding but overlook systematically evaluating navigation agents' spatial perception and reason…

BIP3D: Bridging 2D Images and 3D Perception for Embodied Intelligence

2024-11-22 · CVPR 2025 1 · Xuewu Lin, Tianwei Lin, Lichao Huang, Hongyu Xie 외

In embodied intelligence systems, a key component is 3D perception algorithm, which enables agents to understand their surrounding environments. Previous algorithms primarily rely on point cloud, which, despite offering …

3D visual groundingVisual Grounding

AirGroundBench: Probing Spatial Intelligence in Multimodal Large Models under Heterogeneous Multi-View Embodied Collaboration

2026-06-26 · Haotian Li, Yida Wang, Leyuan Wang, Jinshan Lai 외 arxiv

In recent years, multimodal large language models (MLLMs) have shown strong potential for embodied intelligence, yet their ability to maintain geometrically consistent spatial understanding across heterogeneous views rem…

Vision-Language NavigationVisual Question Answering