paper-with-me

Papers

DriveVLM: The Convergence of Autonomous Driving and Large Vision-Language Models

2024-02-19 · Xiaoyu Tian, Junru Gu, Bailin Li, Yicheng Liu, Yang Wang, Zhiyong Zhao, Kun Zhan, Peng Jia, Xianpeng Lang, Hang Zhao

A primary hurdle of autonomous driving in urban environments is understanding complex and long-tail scenarios, such as challenging road conditions and delicate human behaviors. We introduce DriveVLM, an autonomous driving system leveraging Vision-Language Models (VLMs) for enhanced scene understanding and planning capabilities. DriveVLM integrates a unique combination of reasoning modules for scene description, scene analysis, and hierarchical planning. Furthermore, recognizing the limitations of VLMs in spatial reasoning and heavy computational requirements, we propose DriveVLM-Dual, a hybrid system that synergizes the strengths of DriveVLM with the traditional autonomous driving pipeline. Experiments on both the nuScenes dataset and our SUP-AD dataset demonstrate the efficacy of DriveVLM and DriveVLM-Dual in handling complex and unpredictable driving conditions. Finally, we deploy the DriveVLM-Dual on a production vehicle, verifying it is effective in real-world autonomous driving environments.

📄 PDF Abstract BibTeX arXiv:2402.12289

Code (0)

등록된 구현이 없습니다.

Tasks

Autonomous DrivingScene UnderstandingSpatial Reasoning

Similar Papers 제목 키워드 기반

AutoTrust: Benchmarking Trustworthiness in Large Vision Language Models for Autonomous Driving

2024-12-19 · Shuo Xing, Hongyuan Hua, Xiangbo Gao, Shenzhe Zhu 외

Recent advancements in large vision language models (VLMs) tailored for autonomous driving (AD) have shown strong scene understanding and reasoning capabilities, making them undeniable candidates for end-to-end driving s…

Autonomous DrivingBenchmarkingFairnessQuestion Answering+2

NaviDriveVLM: Decoupling High-Level Reasoning and Motion Planning for Autonomous Driving

2026-03-09 · Ximeng Tao, Pardis Taghavi, Dimitar Filev, Reza Langari 외 arxiv

Vision-language models (VLMs) have emerged as a promising direction for end-to-end autonomous driving (AD) by jointly modeling visual observations, driving context, and language-based reasoning. However, existing VLM-bas…

Autonomous DrivingMotion Planning

DriveVLM-RL: Neuroscience-Inspired Reinforcement Learning with Vision-Language Models for Safe and Deployable Autonomous Driving

2026-03-18 · Zilin Huang, Zihao Sheng, Zhengyang Wan, Yansong Qu 외 arxiv

Traditional reinforcement learning (RL) methods rely on manually engineered rewards or sparse collision signals, which fail to capture the rich contextual understanding required for safe driving and make unsafe explorati…

Reinforcement LearningCollision AvoidanceAutonomous Driving

RoboDriveVLM: A Novel Benchmark and Baseline towards Robust Vision-Language Models for Autonomous Driving

2025-12-01 · Dacheng Liao, Mengshi Qi, Peng Shu, Zhining Zhang 외 arxiv

Current Vision-Language Model (VLM)-based end-to-end autonomous driving systems often leverage large language models to generate driving decisions directly based on their understanding of the current scene. However, such…

Knowledge DistillationTrajectory PredictionTest-time AdaptationAutonomous Driving

CritiqueDriveVLM: From Verifier-Guided Reinforcement Learning to Latent Thought Distillation for Autonomous Driving

2026-07-05 · Zhaohong Liu, Hao Ye, Xianlin Zhang, Mengshi Qi arxiv

End-to-end Vision-Language Models (VLMs) show immense potential in autonomous driving. However, standard Supervised Fine-Tuning (SFT) often suffers from reasoning hallucinations and conservative biases. While traditional…

Reinforcement LearningAutonomous Driving