paper-with-me

홈 › Papers

How Good are Foundation Models in Step-by-Step Embodied Reasoning?

2025-09-18 · Dinura Dissanayake, Ahmed Heakl, Omkar Thawakar, Noor Ahsan, Ritesh Thawkar, Ketan More, Jean Lahoud, Rao Anwer, Hisham Cholakkal, Ivan Laptev, Fahad Shahbaz Khan, Salman Khan arxiv

Embodied agents operating in the physical world must make decisions that are not only effective but also safe, spatially coherent, and grounded in context. While recent advances in large multimodal models (LMMs) have shown promising capabilities in visual understanding and language generation, their ability to perform structured reasoning for real-world embodied tasks remains underexplored. In this work, we aim to understand how well foundation models can perform step-by-step reasoning in embodied environments. To this end, we propose the Foundation Model Embodied Reasoning (FoMER) benchmark, designed to evaluate the reasoning capabilities of LMMs in complex embodied decision-making scenarios. Our benchmark spans a diverse set of tasks that require agents to interpret multimodal observations, reason about physical constraints and safety, and generate valid next actions in natural language. We present (i) a large-scale, curated suite of embodied reasoning tasks, (ii) a novel evaluation framework that disentangles perceptual grounding from action reasoning, and (iii) empirical analysis of several leading LMMs under this setting. Our benchmark includes over 1.1k samples with detailed step-by-step reasoning across 10 tasks and 8 embodiments, covering three different robot types. Our results highlight both the potential and current limitations of LMMs in embodied reasoning, pointing towards key challenges and opportunities for future research in robot intelligence. Our data and code will be made publicly available.

📄 PDF Abstract BibTeX arXiv:2509.15293

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

RxBrain: Embodied Cognition Foundation Model with Joint Language-Visual Reasoning and Imagination

2026-07-15 · Haotian Liang, Mingkang Chen, Yufei Huang, Yuchun Guo 외 hf

Embodied cognition requires agents to connect high-level task reasoning with the physical states to be achieved. We introduce Hy-Embodied-RxBrain, an embodied cognition foundation model with joint language-visual reasoni…

Scene UnderstandingVisual ReasoningDecision Making

Nav-R1: Reasoning and Navigation in Embodied Scenes

2025-09-13 · Qingxiang Liu, Ting Huang, Zeyu Zhang, Hao Tang arxiv

Embodied navigation requires agents to integrate perception, reasoning, and action for robust interaction in complex 3D environments. Existing approaches often suffer from incoherent and unstable reasoning traces that hi…

Reinforcement Learning

Towards Embodied Cognition in Robots via Spatially Grounded Synthetic Worlds

2025-05-20 · Joel Currie, Gioele Migno, Enrico Piacenti, Maria Elena Giannaccini 외

We present a conceptual framework for training Vision-Language Models (VLMs) to perform Visual Perspective Taking (VPT), a core capability for embodied cognition essential for Human-Robot Interaction (HRI). As a first st…

Spatial Reasoning

Vlaser: Vision-Language-Action Model with Synergistic Embodied Reasoning

2025-10-13 · Ganlin Yang, Tianyi Zhang, Haoran Hao, Weiyun Wang 외 arxiv

While significant research has focused on developing embodied reasoning capabilities using Vision-Language Models (VLMs) or integrating advanced VLMs into Vision-Language-Action (VLA) models for end-to-end robot control,…

Spatial Reasoning

RoboPIN: Grounded Embodied Reasoning via Pinned Chain-of-Thought

2026-06-14 · Yaoting Huang, Yifu Yuan, Linqi Han, Chengwen Li 외 arxiv

Embodied reasoning requires models to perceive task-relevant objects and spaces in physical environments and maintain consistent visual grounding throughout multi-step reasoning. However, current vision-language models r…

Spatial ReasoningVisual Grounding