paper-with-me

홈 › Papers

DriveDreamer-2: LLM-Enhanced World Models for Diverse Driving Video Generation

2024-03-11 · Guosheng Zhao, XiaoFeng Wang, Zheng Zhu, Xinze Chen, Guan Huang, Xiaoyi Bao, Xingang Wang

World models have demonstrated superiority in autonomous driving, particularly in the generation of multi-view driving videos. However, significant challenges still exist in generating customized driving videos. In this paper, we propose DriveDreamer-2, which builds upon the framework of DriveDreamer and incorporates a Large Language Model (LLM) to generate user-defined driving videos. Specifically, an LLM interface is initially incorporated to convert a user's query into agent trajectories. Subsequently, a HDMap, adhering to traffic regulations, is generated based on the trajectories. Ultimately, we propose the Unified Multi-View Model to enhance temporal and spatial coherence in the generated driving videos. DriveDreamer-2 is the first world model to generate customized driving videos, it can generate uncommon driving videos (e.g., vehicles abruptly cut in) in a user-friendly manner. Besides, experimental results demonstrate that the generated videos enhance the training of driving perception methods (e.g., 3D detection and tracking). Furthermore, video generation quality of DriveDreamer-2 surpasses other state-of-the-art methods, showcasing FID and FVD scores of 11.2 and 55.7, representing relative improvements of 30% and 50%.

📄 PDF Abstract BibTeX arXiv:2403.06845

Code (1)

f1yfisher/drivedreamer2 pytorch

Tasks

Autonomous DrivingLanguage ModelingLanguage ModellingLarge Language ModelVideo Generation

Similar Papers 제목 키워드 기반

DriveDreamer4D: World Models Are Effective Data Machines for 4D Driving Scene Representation

2024-10-17 · CVPR 2025 1 · Guosheng Zhao, Chaojun Ni, XiaoFeng Wang, Zheng Zhu 외

Closed-loop simulation is essential for advancing end-to-end autonomous driving systems. Contemporary sensor simulation methods, such as NeRF and 3DGS, rely predominantly on conditions closely aligned with training data …

3DGS4D reconstructionAutonomous DrivingNeRF+1

DriveDreamer: Towards Real-world-driven World Models for Autonomous Driving

2023-09-18 · XiaoFeng Wang, Zheng Zhu, Guan Huang, Xinze Chen 외

World models, especially in autonomous driving, are trending and drawing extensive attention due to their capacity for comprehending driving environments. The established world model holds immense potential for the gener…

Autonomous DrivingVideo Generation

UniDriveDreamer: A Single-Stage Multimodal World Model for Autonomous Driving

2026-02-02 · Guosheng Zhao, Yaozeng Wang, Xiaofeng Wang, Zheng Zhu 외 arxiv

World models have demonstrated significant promise for data synthesis in autonomous driving. However, existing methods predominantly concentrate on single-modality generation, typically focusing on either multi-camera vi…

Autonomous Driving

DriveDreamer-Policy: A Geometry-Grounded World-Action Model for Unified Generation and Planning

2026-04-02 · Yang Zhou, Xiaofeng Wang, Hao Shao, Letian Wang 외 arxiv

Recently, world-action models (WAM) have emerged to bridge vision-language-action (VLA) models and world models, unifying their reasoning and instruction-following capabilities and spatio-temporal world modeling. However…

Video GenerationMotion Planning

WorldSplat: Gaussian-Centric Feed-Forward 4D Scene Generation for Autonomous Driving

2025-09-27 · Ziyue Zhu, Zhanqian Wu, Zhenxin Zhu, Lijun Zhou 외 arxiv

Recent advances in driving-scene generation and reconstruction have demonstrated significant potential for enhancing autonomous driving systems by producing scalable and controllable training data. Existing generation me…

Autonomous DrivingScene Generation