paper-with-me

홈 › Papers

Drive-P2D: A Progressive Perception-to-Decision Benchmark for VLMs in Autonomous Driving

2026-01-21 · Zecong Tang, Zixu Wang, Yifei Wang, Weitong Lian, Tianjian Gao, Haoran Li, Tengju Ru, Lingyi Meng, Zhejun Cui, Yichen Zhu, Qi Kang, Kaixuan Wang, Yu Zhang arxiv

Autonomous driving requires reliable perception and safe decision-making in complex scenarios. Recent vision-language models (VLMs) demonstrate reasoning and generalization abilities, opening new possibilities for autonomous driving; however, existing benchmarks often evaluate perception and decision-making separately, limit failure analysis with choice-only formats, or introduce evaluation bias through LLM-scored long-form outputs. To address these issues, we present Drive-P2D, a progressive perception-to-decision benchmark with 6,650 questions across Object, Scene, and Decision levels. Drive-P2D adopts a separated reasoning-and-answer protocol: final answers are scored objectively, while reasoning is analyzed to identify error modes exposed along the progressive perception-to-decision chain. We evaluate mainstream VLMs across all and high-risk scenarios, and further characterize the perception-to-decision capability boundary through correlation analysis and similar-scene robustness testing. Reasoning further exposes failure modes such as logical reasoning errors and semantic feature omissions, and we train a lightweight analyzer model to automate large-scale error-mode annotation of reasoning. Together, these designs provide practical insights for building safer and more reliable VLMs for real-world autonomous driving.

📄 PDF Abstract BibTeX arXiv:2601.14702

Code (0)

등록된 구현이 없습니다.

Tasks

Autonomous DrivingLogical Reasoning

Similar Papers 제목 키워드 기반

$AutoDrive\text{-}P^3$: Unified Chain of Perception-Prediction-Planning Thought via Reinforcement Fine-Tuning

2026-03-30 · Yuqi Ye, Zijian Zhang, Junhong Lin, Shangkun Sun 외 arxiv

Vision-language models (VLMs) are increasingly being adopted for end-to-end autonomous driving systems due to their exceptional performance in handling long-tail scenarios. However, current VLM-based approaches suffer fr…

Hierarchical Reinforcement LearningAutonomous Driving

GUI-Eyes: Tool-Augmented Perception for Visual Grounding in GUI Agents

2026-01-14 · Chen Chen, Jiawei Shao, Dakuan Lu, Haoyi Hu 외 arxiv

Recent advances in vision-language models (VLMs) and reinforcement learning (RL) have driven progress in GUI automation. However, most existing methods rely on static, one-shot visual inputs and passive perception, lacki…

Reinforcement LearningVisual Grounding

SpatialCLI: Learning to Reason With Spatial Tools, Then Without Them

2026-07-30 · Yang Zhou, Zixuan Huang, Sunzhu Li, Zhuo Yang 외 arxiv

Vision-language models (VLMs) are increasingly used in embodied agents to interpret visual inputs, reason about spatial relationships, and make task-level decisions based on that reasoning. However, a fundamental capabil…

WorkDrive: Roadwork Chain of Causation for Autonomous Driving

2026-07-16 · Tianyi Jiang, Wen Zhang, Sihan Yang, Ming Lu 외 arxiv

Autonomous driving vision-language models (VLMs) struggle in roadwork zones, where familiar visual cues such as lane markings and permanent signs are altered or absent, and temporary devices such as cones and barriers re…

Reinforcement LearningTrajectory PredictionAutonomous Driving

PlanBench-V: A Spatial Planning Map Benchmark for Vision-Language Models

2026-06-04 · Minxin Chen, He Zhu, Junyou Su, Wen Wang 외 arxiv

Spatial planning maps are central to territorial governance, translating planning objectives, regulations, and spatial strategies into visual forms for decision-making, public communication, and institutional coordinatio…

Multimodal ReasoningSpatial Reasoning