paper-with-me

홈 › Papers

Bench2Drive-VL: Benchmarks for Closed-Loop Autonomous Driving with Vision-Language Models

2026-04-01 · Xiaosong Jia, Yuqian Shao, Zhenjie Yang, Qifeng Li, Zhiyuan Zhang, Junchi Yan arxiv

With the rise of vision-language models (VLM), their application for autonomous driving (VLM4AD) has gained significant attention. Meanwhile, in autonomous driving, closed-loop evaluation has become widely recognized as a more reliable validation method than open-loop evaluation, as it can evaluate the performance of the model under cumulative errors and out-of-distribution inputs. However, existing VLM4AD benchmarks evaluate the model`s scene understanding ability under open-loop, i.e., via static question-answer (QA) dataset. This kind of evaluation fails to assess the VLMs performance under out-of-distribution states rarely appeared in the human collected datasets.To this end, we present Bench2Drive-VL, an extension of Bench2Drive that brings closed-loop evaluation to VLM-based driving, which introduces: (1) DriveCommenter, a closed-loop generator that automatically generates diverse, behavior-grounded question-answer pairs for all driving situations in CARLA,including severe off-route and off-road deviations previously unassessable in simulation. (2) A unified protocol and interface that allows modern VLMs to be directly plugged into the Bench2Drive closed-loop environment to compare with traditional agents. (3) A flexible reasoning and control framework, supporting multi-format visual inputs and configurable graph-based chain-of-thought execution. (4) A complete development ecosystem. Together, these components form a comprehensive closed-loop benchmark for VLM4AD. All codes and annotated datasets are open sourced.

📄 PDF Abstract BibTeX arXiv:2604.01259

Code (0)

등록된 구현이 없습니다.

Tasks

Scene UnderstandingAutonomous Driving

Similar Papers 제목 키워드 기반

X-Driver: Explainable Autonomous Driving with Vision-Language Models

2025-05-08 · Wei Liu, Jiyuan Zhang, Binxiong Zheng, Yufeng Hu 외

End-to-end autonomous driving has advanced significantly, offering benefits such as system simplicity and stronger driving performance in both open-loop and closed-loop settings than conventional pipelines. However, exis…

Autonomous DrivingBench2DriveDecision Making

HiDrive: A Closed-Loop Benchmark for High-Level Autonomous Driving

2026-05-11 · Zhongyu Xia, Guanyu Zhu, Guo Tang, Wenhao Chen 외 arxiv

End-to-end autonomous driving has witnessed rapid progress, yet existing benchmarks are increasingly saturated, with state-of-the-art models achieving near-perfect scores on widely used open-loop and closed-loop benchmar…

Autonomous Driving

Bench2Drive-Robust: Benchmarking Closed-Loop Autonomous Driving under Deployment Perturbations

2026-05-18 · Zhiyuan Zhang, Zhenghao Jin, Yanlun Peng, Xianda Guo 외 arxiv

Robustness is a critical requirement for deploying autonomous driving systems in the real world. Existing robustness benchmarks for autonomous driving have made important progress in studying the effects of image-level c…

Autonomous Driving

DriveE2E: Closed-Loop Benchmark for End-to-End Autonomous Driving through Real-to-Simulation

2025-09-28 · Haibao Yu, Wenxian Yang, Ruiyang Hao, Chuanye Wang 외 arxiv

Closed-loop evaluation is increasingly critical for end-to-end autonomous driving. Current closed-loop benchmarks using the CARLA simulator rely on manually configured traffic scenarios, which can diverge from real-world…

Autonomous Driving

Fail2Drive: Benchmarking Closed-Loop Driving Generalization

2026-04-09 · Simon Gerstenecker, Andreas Geiger, Katrin Renz arxiv

Generalization under distribution shift remains a central bottleneck for closed-loop autonomous driving. Although simulators like CARLA enable safe and scalable testing, existing benchmarks rarely measure true generaliza…

Autonomous Driving