paper-with-me

홈 › Papers

A Survey on Multimodal Large Language Models for Autonomous Driving

2023-11-21 · Can Cui, Yunsheng Ma, Xu Cao, Wenqian Ye, Yang Zhou, Kaizhao Liang, Jintai Chen, Juanwu Lu, Zichong Yang, Kuei-Da Liao, Tianren Gao, Erlong Li, Kun Tang, Zhipeng Cao, Tong Zhou, Ao Liu, Xinrui Yan, Shuqi Mei, Jianguo Cao, Ziran Wang, Chao Zheng

With the emergence of Large Language Models (LLMs) and Vision Foundation Models (VFMs), multimodal AI systems benefiting from large models have the potential to equally perceive the real world, make decisions, and control tools as humans. In recent months, LLMs have shown widespread attention in autonomous driving and map systems. Despite its immense potential, there is still a lack of a comprehensive understanding of key challenges, opportunities, and future endeavors to apply in LLM driving systems. In this paper, we present a systematic investigation in this field. We first introduce the background of Multimodal Large Language Models (MLLMs), the multimodal models development using LLMs, and the history of autonomous driving. Then, we overview existing MLLM tools for driving, transportation, and map systems together with existing datasets and benchmarks. Moreover, we summarized the works in The 1st WACV Workshop on Large Language and Vision Models for Autonomous Driving (LLVM-AD), which is the first workshop of its kind regarding LLMs in autonomous driving. To further promote the development of this field, we also discuss several important problems regarding using MLLMs in autonomous driving systems that need to be solved by both academia and industry.

📄 PDF Abstract BibTeX arXiv:2311.12320

Code (1)

irohxu/awesome-multimodal-llm-autonomous-driving 공식 구현

Tasks

Autonomous Driving

Similar Papers 제목 키워드 기반

A Survey on Vision-Language-Action Models for Autonomous Driving

2025-06-30 · Sicong Jiang, Zilin Huang, Kangan Qian, Ziang Luo 외

The rapid progress of multimodal large language models (MLLM) has paved the way for Vision-Language-Action (VLA) paradigms, which integrate visual perception, natural language understanding, and control within a single p…

Autonomous DrivingAutonomous VehiclesNatural Language UnderstandingSurvey+1

Foundation Models in Autonomous Driving: A Survey on Scenario Generation and Scenario Analysis

2025-06-13 · Yuan Gao, Mattia Piccinini, Yuchen Zhang, Dingrui Wang 외

For autonomous vehicles, safe navigation in complex environments depends on handling a broad range of diverse and rare driving scenarios. Simulation- and scenario-based testing have emerged as key approaches to developme…

Autonomous DrivingAutonomous VehiclesSurvey

XLM for Autonomous Driving Systems: A Comprehensive Review

2024-09-16 · Sonda Fourati, Wael Jaafar, Noura Baccar, Safwan Alfattani

Large Language Models (LLMs) have showcased remarkable proficiency in various information-processing tasks. These tasks span from extracting data and summarizing literature to generating content, predictive modeling, dec…

Autonomous DrivingDecision Making

Large Foundation Models for Trajectory Prediction in Autonomous Driving: A Comprehensive Survey

2025-09-11 · Wei Dai, Shengen Wu, Wei Wu, Zhenhao Wang 외 arxiv

Trajectory prediction serves as a critical functionality in autonomous driving, enabling the anticipation of future motion paths for traffic participants such as vehicles and pedestrians, which is essential for driving s…

Trajectory PredictionAutonomous Driving

Progressive Bird's Eye View Perception for Safety-Critical Autonomous Driving: A Comprehensive Survey

2025-08-11 · Yan Gong, Naibang Wang, Jianli Lu, Xinyu Zhang 외 arxiv

Bird's-Eye-View (BEV) perception has become a foundational paradigm in autonomous driving, enabling unified spatial representations that support robust multi-sensor fusion and multi-agent collaboration. As autonomous veh…

Autonomous VehiclesAutonomous Driving