paper-with-me

홈 › Papers

MUVO: A Multimodal World Model with Spatial Representations for Autonomous Driving

2023-11-20 · Daniel Bogdoll, Yitian Yang, Tim Joseph, J. Marius Zöllner

Learning unsupervised world models for autonomous driving has the potential to improve the reasoning capabilities of today's systems dramatically. However, most work neglects the physical attributes of the world and focuses on sensor data alone. We propose MUVO, a MUltimodal World Model with spatial VOxel representations, to address this challenge. We utilize raw camera and lidar data to learn a sensor-agnostic geometric representation of the world. We demonstrate multimodal future predictions and show that our spatial representation improves the prediction quality of both camera images and lidar point clouds.

📄 PDF Abstract BibTeX arXiv:2311.11762

Code (1)

fzi-forschungszentrum-informatik/muvo 공식 구현 pytorch

Tasks

Autonomous Driving

Similar Papers 제목 키워드 기반

MUVOD: A Novel Multi-view Video Object Segmentation Dataset and A Benchmark for 3D Segmentation

2025-07-10 · Bangning Wei, Joshua Maraval, Meriem Outtas, Kidiyo Kpalma 외

The application of methods based on Neural Radiance Fields (NeRF) and 3D Gaussian Splatting (3D GS) have steadily gained popularity in the field of 3D object segmentation in static scenes. These approaches demonstrate ef…

NeRFObjectScene UnderstandingSegmentation+4

Think Before You Drive: World Model-Inspired Multimodal Grounding for Autonomous Vehicles

2025-12-03 · Haicheng Liao, Huanming Shen, Bonan Wang, Yongkang Li 외 arxiv

Interpreting natural-language commands to localize target objects is critical for autonomous driving (AD). Existing visual grounding (VG) methods for autonomous vehicles (AVs) typically struggle with ambiguous, context-d…

Autonomous VehiclesAutonomous DrivingVisual Grounding

OccVLA: Vision-Language-Action Model with Implicit 3D Occupancy Supervision

2025-09-06 · Ruixun Liu, Lingyu Kong, Derun Li, Hang Zhao arxiv

Multimodal large language models (MLLMs) have shown strong vision-language reasoning abilities but still lack robust 3D spatial understanding, which is critical for autonomous driving. This limitation stems from two key …

Multimodal ReasoningTrajectory PlanningAutonomous Driving

Progressive Bird's Eye View Perception for Safety-Critical Autonomous Driving: A Comprehensive Survey

2025-08-11 · Yan Gong, Naibang Wang, Jianli Lu, Xinyu Zhang 외 arxiv

Bird's-Eye-View (BEV) perception has become a foundational paradigm in autonomous driving, enabling unified spatial representations that support robust multi-sensor fusion and multi-agent collaboration. As autonomous veh…

Autonomous VehiclesAutonomous Driving

Latent-WAM: Latent World Action Modeling for End-to-End Autonomous Driving

2026-03-25 · Linbo Wang, Yupeng Zheng, Qiang Chen, Shiwei Li 외 arxiv

We introduce Latent-WAM, an efficient end-to-end autonomous driving framework that achieves strong trajectory planning through spatially-aware and dynamics-informed latent world representations. Existing world-model-base…

Trajectory PlanningAutonomous Driving