paper-with-me

홈 › Papers

SimWorld-Robotics: Synthesizing Photorealistic and Dynamic Urban Environments for Multimodal Robot Navigation and Collaboration

2025-12-10 · Yan Zhuang, Jiawei Ren, Xiaokang Ye, Jianzhi Shen, Ruixuan Zhang, Tianai Yue, Muhammad Faayez, Xuhong He, Ziqiao Ma, Lianhui Qin, Zhiting Hu, Tianmin Shu arxiv

Recent advances in foundation models have shown promising results in developing generalist robotics that can perform diverse tasks in open-ended scenarios given multimodal inputs. However, current work has been mainly focused on indoor, household scenarios. In this work, we present SimWorld-Robotics~(SWR), a simulation platform for embodied AI in large-scale, photorealistic urban environments. Built on Unreal Engine 5, SWR procedurally generates unlimited photorealistic urban scenes populated with dynamic elements such as pedestrians and traffic systems, surpassing prior urban simulations in realism, complexity, and scalability. It also supports multi-robot control and communication. With these key features, we build two challenging robot benchmarks: (1) a multimodal instruction-following task, where a robot must follow vision-language navigation instructions to reach a destination in the presence of pedestrians and traffic; and (2) a multi-agent search task, where two robots must communicate to cooperatively locate and meet each other. Unlike existing benchmarks, these two new benchmarks comprehensively evaluate a wide range of critical robot capacities in realistic scenarios, including (1) multimodal instructions grounding, (2) 3D spatial reasoning in large environments, (3) safe, long-range navigation with people and traffic, (4) multi-robot collaboration, and (5) grounded communication. Our experimental results demonstrate that state-of-the-art models, including vision-language models (VLMs), struggle with our tasks, lacking robust perception, reasoning, and planning abilities necessary for urban environments.

📄 PDF Abstract BibTeX arXiv:2512.10046

Code (0)

등록된 구현이 없습니다.

Tasks

Vision-Language NavigationSpatial ReasoningRobot Navigation

Similar Papers 제목 키워드 기반

SimWorld: An Open-ended Realistic Simulator for Autonomous Agents in Physical and Social Worlds

2025-11-30 · Jiawei Ren, Yan Zhuang, Xiaokang Ye, Lingjun Mao 외 arxiv

While LLM/VLM-powered AI agents have advanced rapidly in math, coding, and computer use, their applications in complex physical and social environments remain challenging. Building agents that can survive and thrive in t…

SimWorlds: A Multi-Agent System for Dynamic 3D Scene Creation

2026-07-02 · Chunjiang Liu, Xiaoyuan Wang, Haoyu Chen, Yizhou Zhao 외 arxiv

LLM agents are increasingly used to translate natural language into 3D scenes in a procedural way, but existing systems focus on static output. Dynamic 4D scenes from text alone, in which liquids flow, particles emit, ri…

Video Generation

Skyfall-GS: Synthesizing Immersive 3D Urban Scenes from Satellite Imagery

2025-10-17 · Jie-Ying Lee, Yi-Ruei Liu, Shr-Ruei Tsai, Wei-Cheng Chang 외 arxiv

Synthesizing large-scale, explorable, and geometrically accurate 3D urban scenes is a challenging yet valuable task for immersive and embodied applications. The challenge lies in the lack of large-scale and high-quality …

SimWorld: A Unified Benchmark for Simulator-Conditioned Scene Generation via World Model

2025-03-18 · Xinqing Li, Ruiqi Song, Qingyu Xie, Ye Wu 외

With the rapid advancement of autonomous driving technology, a lack of data has become a major obstacle to enhancing perception model accuracy. Researchers are now exploring controllable data generation using world model…

Autonomous DrivingImage GenerationScene Generation

Video Generation Models in Robotics -- Applications, Research Challenges, Future Directions

2026-01-12 · Zhiting Mei, Tenny Yin, Ola Shorinwa, Apurva Badithela 외 arxiv

Video generation models have emerged as high-fidelity models of the physical world, capable of synthesizing high-quality videos capturing fine-grained interactions between agents and their environments conditioned on mul…

Reinforcement LearningInstruction FollowingVideo Generation