paper-with-me

홈 › Papers

GRUtopia: Dream General Robots in a City at Scale

2024-07-15 · Hanqing Wang, Jiahe Chen, Wensi Huang, Qingwei Ben, Tai Wang, Boyu Mi, Tao Huang, Siheng Zhao, Yilun Chen, Sizhe Yang, Peizhou Cao, Wenye Yu, Zichao Ye, Jialun Li, Junfeng Long, ZiRui Wang, Huiling Wang, Ying Zhao, Zhongying Tu, Yu Qiao, Dahua Lin, Jiangmiao Pang

Recent works have been exploring the scaling laws in the field of Embodied AI. Given the prohibitive costs of collecting real-world data, we believe the Simulation-to-Real (Sim2Real) paradigm is a crucial step for scaling the learning of embodied models. This paper introduces project GRUtopia, the first simulated interactive 3D society designed for various robots. It features several advancements: (a) The scene dataset, GRScenes, includes 100k interactive, finely annotated scenes, which can be freely combined into city-scale environments. In contrast to previous works mainly focusing on home, GRScenes covers 89 diverse scene categories, bridging the gap of service-oriented environments where general robots would be initially deployed. (b) GRResidents, a Large Language Model (LLM) driven Non-Player Character (NPC) system that is responsible for social interaction, task generation, and task assignment, thus simulating social scenarios for embodied AI applications. (c) The benchmark, GRBench, supports various robots but focuses on legged robots as primary agents and poses moderately challenging tasks involving Object Loco-Navigation, Social Loco-Navigation, and Loco-Manipulation. We hope that this work can alleviate the scarcity of high-quality data in this field and provide a more comprehensive assessment of Embodied AI research. The project is available at https://github.com/OpenRobotLab/GRUtopia.

📄 PDF Abstract BibTeX arXiv:2407.10943

Code (1)

openrobotlab/grutopia 공식 구현

Tasks

Language ModellingLarge Language Model

Similar Papers 제목 키워드 기반

Grasp as You Dream: Imitating Functional Grasping from Generated Human Demonstrations

2026-04-08 · Chao Tang, Jiacheng Xu, Haofei Lu, Bolin Zou 외 arxiv

Building generalist robots capable of performing functional grasping in everyday, open-world environments remains a significant challenge due to the vast diversity of objects and tasks. Existing methods are either constr…

Video Generation

Morphology-Consistent Humanoid Interaction through Robot-Centric Video Synthesis

2026-03-20 · Weisheng Xu, Jian Li, Yi Gu, Bin Yang 외 arxiv

Equipping humanoid robots with versatile interaction skills typically requires either extensive policy training or explicit human-to-robot motion retargeting. However, learning-based policies face prohibitive data collec…

Video GenerationPose Estimation

DreamToNav: Generalizable Navigation for Robots via Generative Video Planning

2026-03-06 · Valerii Serpiva, Jeffrin Sam, Chidera Simon, Hajira Amjad 외 arxiv

We present DreamToNav, a novel autonomous robot framework that uses generative video models to enable intuitive, human-in-the-loop control. Instead of relying on rigid waypoint navigation, users provide natural language …

Video GenerationPose Estimation

QuaDreamer: Controllable Panoramic Video Generation for Quadruped Robots

2025-08-04 · Sheng Wu, Fei Teng, Hao Shi, Qi Jiang 외 arxiv

Panoramic cameras, capturing comprehensive 360-degree environmental data, are suitable for quadruped robots in surrounding perception and interaction with complex environments. However, the scarcity of high-quality panor…

Multi-Object TrackingVideo Generation

DayDreamer: World Models for Physical Robot Learning

2022-06-28 · Philipp Wu, Alejandro Escontrela, Danijar Hafner, Ken Goldberg 외

To solve tasks in complex environments, robots need to learn from experience. Deep reinforcement learning is a common approach to robot learning but requires a large amount of trial and error to learn, limiting its deplo…

Deep Reinforcement LearningNavigatereinforcement-learningReinforcement Learning (RL)