paper-with-me

홈 › Papers

DriveLaW:Unifying Planning and Video Generation in a Latent Driving World

2025-12-29 · Tianze Xia, Yongkang Li, Lijun Zhou, Jingfeng Yao, Kaixin Xiong, Haiyang Sun, Bing Wang, Kun Ma, Guang Chen, Hangjun Ye, Wenyu Liu, Xinggang Wang arxiv

World models have become crucial for autonomous driving, as they learn how scenarios evolve over time to address the long-tail challenges of the real world. However, current approaches relegate world models to limited roles: they operate within ostensibly unified architectures that still keep world prediction and motion planning as decoupled processes. To bridge this gap, we propose DriveLaW, a novel paradigm that unifies video generation and motion planning. By directly injecting the latent representation from its video generator into the planner, DriveLaW ensures inherent consistency between high-fidelity future generation and reliable trajectory planning. Specifically, DriveLaW consists of two core components: DriveLaW-Video, our powerful world model that generates high-fidelity forecasting with expressive latent representations, and DriveLaW-Act, a diffusion planner that generates consistent and reliable trajectories from the latent of DriveLaW-Video, with both components optimized by a three-stage progressive training strategy. The power of our unified paradigm is demonstrated by new state-of-the-art results across both tasks. DriveLaW not only advances video prediction significantly, surpassing best-performing work by 33.3% in FID and 1.8% in FVD, but also achieves a new record on the NAVSIM planning benchmark.

📄 PDF Abstract BibTeX arXiv:2512.23421

Code (0)

등록된 구현이 없습니다.

Tasks

Trajectory PlanningAutonomous DrivingVideo GenerationVideo Prediction

Similar Papers 제목 키워드 기반

Bridging Scene Generation and Planning: Driving with World Model via Unifying Vision and Motion Representation

2026-03-16 · Xingtai Gui, Meijie Zhang, Tianyi Yan, Wencheng Han 외 arxiv

End-to-end autonomous driving aims to generate safe and plausible planning policies from raw sensor input. Driving world models have shown great potential in learning rich representations by predicting the future evoluti…

Autonomous DrivingScene GenerationVideo GenerationMotion Planning

DriveDreamer-Policy: A Geometry-Grounded World-Action Model for Unified Generation and Planning

2026-04-02 · Yang Zhou, Xiaofeng Wang, Hao Shao, Letian Wang 외 arxiv

Recently, world-action models (WAM) have emerged to bridge vision-language-action (VLA) models and world models, unifying their reasoning and instruction-following capabilities and spatio-temporal world modeling. However…

Video GenerationMotion Planning

Image Generation as a Visual Planner for Robotic Manipulation

2025-11-29 · Ye Pang arxiv

Generating realistic robotic manipulation videos is an important step toward unifying perception, planning, and action in embodied agents. While existing video diffusion models require large domain-specific datasets and …

Image Generation

DrivingGPT: Unifying Driving World Modeling and Planning with Multi-modal Autoregressive Transformers

2024-12-24 · Yuntao Chen, Yuqi Wang, Zhaoxiang Zhang

World model-based searching and planning are widely recognized as a promising path toward human-level physical intelligence. However, current driving world models primarily rely on video diffusion models, which specializ…

NavSimTrajectory PlanningVideo Generation

UniUGP: Unifying Understanding, Generation, and Planing For End-to-end Autonomous Driving

2025-12-10 · Hao Lu, Ziyang Liu, Guangfeng Jiang, Yuanfei Luo 외 arxiv

Autonomous driving (AD) systems struggle in long-tail scenarios due to limited world knowledge and weak visual dynamic modeling. Existing vision-language-action (VLA)-based methods cannot leverage unlabeled videos for vi…

Trajectory PlanningAutonomous DrivingVideo Generation