paper-with-me

Papers

World Model for Robot Learning: A Comprehensive Survey

2026-04-30 · Bohan Hou, Gen Li, Jindou Jia, Tuo An, Xinying Guo, Sicong Leng, Haoran Geng, Yanjie Ze, Tatsuya Harada, Philip Torr, Oier Mees, Marc Pollefeys, Zhuang Liu, Jiajun Wu, Pieter Abbeel, Jitendra Malik, Yilun Du, Jianfei Yang arxiv

World models, which are predictive representations of how environments evolve under actions, have become a central component of robot learning. They support policy learning, planning, simulation, evaluation, data generation, and have advanced rapidly with the rise of foundation models and large-scale video generation. However, the literature remains fragmented across architectures, functional roles, and embodied application domains. To address this gap, we present a comprehensive review of world models from a robot-learning perspective. We examine how world models are coupled with robot policies, how they serve as learned simulators for reinforcement learning and evaluation, and how robotic video world models have progressed from imagination-based generation to controllable, structured, and foundation-scale formulations. We further connect these ideas to navigation and autonomous driving, and summarize representative datasets, benchmarks, and evaluation protocols. Overall, this survey systematically reviews the rapidly growing literature on world models for robot learning, clarifies key paradigms and applications, and highlights major challenges and future directions for predictive modeling in embodied agents. To facilitate continued access to newly emerging works, benchmarks, and resources, we will maintain and regularly update the accompanying GitHub repository alongside this survey.

📄 PDF Abstract BibTeX arXiv:2605.00080

Code (0)

등록된 구현이 없습니다.

Tasks

Reinforcement LearningAutonomous DrivingVideo Generation

Similar Papers 제목 키워드 기반

Vision-Language-Action Models for Robotics: A Review Towards Real-World Applications

2025-10-08 · Kento Kawaharazuka, Jihoon Oh, Jun Yamada, Ingmar Posner 외 arxiv

Amid growing efforts to leverage advances in large language models (LLMs) and vision-language models (VLMs) for robotics, Vision-Language-Action (VLA) models have recently gained significant attention. By unifying vision…

Data Augmentation

A Comprehensive Survey on Physical Risk Control in the Era of Foundation Model-enabled Robotics

2025-05-19 · Takeshi Kojima, Yaonan Zhu, Yusuke Iwasawa, Toshinori Kitamura 외

Recent Foundation Model-enabled robotics (FMRs) display greatly improved general-purpose skills, enabling more adaptable automation than conventional robotics. Their ability to handle diverse tasks thus creates new oppor…

Survey

A Survey: Learning Embodied Intelligence from Physical Simulators and World Models

2025-07-01 · Xiaoxiao Long, Qingrui Zhao, Kaiwen Zhang, Zihao Zhang 외 arxiv

The pursuit of artificial general intelligence (AGI) has placed embodied intelligence at the forefront of robotics research. Embodied intelligence focuses on agents capable of perceiving, reasoning, and acting within the…

State-of-the-art in Robot Learning for Multi-Robot Collaboration: A Comprehensive Survey

2024-08-03 · Bin Wu, C Steve Suh

With the continuous breakthroughs in core technology, the dawn of large-scale integration of robotic systems into daily human life is on the horizon. Multi-robot systems (MRS) built on this foundation are undergoing dras…

Large Language Models for Multi-Robot Systems: A Survey

2025-02-06 · Peihan Li, Zijian An, Shams Abrar, Lifeng Zhou

The rapid advancement of Large Language Models (LLMs) has opened new possibilities in Multi-Robot Systems (MRS), enabling enhanced communication, task planning, and human-robot interaction. Unlike traditional single-robo…

Action GenerationBenchmarkingHallucinationMathematical Reasoning+3