paper-with-me

홈 › Papers

The Essential Role of Causality in Foundation World Models for Embodied AI

2024-02-06 · Tarun Gupta, Wenbo Gong, Chao Ma, Nick Pawlowski, Agrin Hilmkil, Meyer Scetbon, Marc Rigter, Ade Famoti, Ashley Juan Llorens, Jianfeng Gao, Stefan Bauer, Danica Kragic, Bernhard Schölkopf, Cheng Zhang

Recent advances in foundation models, especially in large multi-modal models and conversational agents, have ignited interest in the potential of generally capable embodied agents. Such agents will require the ability to perform new tasks in many different real-world environments. However, current foundation models fail to accurately model physical interactions and are therefore insufficient for Embodied AI. The study of causality lends itself to the construction of veridical world models, which are crucial for accurately predicting the outcomes of possible interactions. This paper focuses on the prospects of building foundation world models for the upcoming generation of embodied agents and presents a novel viewpoint on the significance of causality within these. We posit that integrating causal considerations is vital to facilitating meaningful physical interactions with the world. Finally, we demystify misconceptions about causality in this context and present our outlook for future research.

📄 PDF Abstract BibTeX arXiv:2402.06665

Code (0)

등록된 구현이 없습니다.

Tasks

Misconceptions

Similar Papers 제목 키워드 기반

Aligning Cyber Space with Physical World: A Comprehensive Survey on Embodied AI

2024-07-09 · Yang Liu, Weixing Chen, Yongjie Bai, Xiaodan Liang 외

Embodied Artificial Intelligence (Embodied AI) is crucial for achieving Artificial General Intelligence (AGI) and serves as a foundation for various applications that bridge cyberspace and the physical world. Recently, t…

Survey

Embodied AI with Foundation Models for Mobile Service Robots: A Systematic Review

2025-05-26 · Matthew Lisondra, Beno Benhabib, Goldie Nejat

Rapid advancements in foundation models, including Large Language Models, Vision-Language Models, Multimodal Large Language Models, and Vision-Language-Action Models have opened new avenues for embodied AI in mobile serv…

Decision Making Under UncertaintySensor FusionVision-Language-Action

HY-Embodied-0.5: Embodied Foundation Models for Real-World Agents

2026-04-08 · Tencent Robotics X, HY Vision Team, :, Xumin Yu 외 arxiv

We introduce HY-Embodied-0.5, a family of foundation models specifically designed for real-world embodied agents. To bridge the gap between general Vision-Language Models (VLMs) and the demands of embodied agents, our mo…

Spatial Reasoning

RxBrain: Embodied Cognition Foundation Model with Joint Language-Visual Reasoning and Imagination

2026-07-15 · Haotian Liang, Mingkang Chen, Yufei Huang, Yuchun Guo 외 hf

Embodied cognition requires agents to connect high-level task reasoning with the physical states to be achieved. We introduce Hy-Embodied-RxBrain, an embodied cognition foundation model with joint language-visual reasoni…

Scene UnderstandingVisual ReasoningDecision Making

Towards Embodied Cognition in Robots via Spatially Grounded Synthetic Worlds

2025-05-20 · Joel Currie, Gioele Migno, Enrico Piacenti, Maria Elena Giannaccini 외

We present a conceptual framework for training Vision-Language Models (VLMs) to perform Visual Perspective Taking (VPT), a core capability for embodied cognition essential for Human-Robot Interaction (HRI). As a first st…

Spatial Reasoning