paper-with-me

홈 › Papers

Scalable Offline Metrics for Autonomous Driving

2025-10-09 · Animikh Aich, Adwait Kulkarni, Eshed Ohn-Bar arxiv

Real-world evaluation of perception-based planning models for robotic systems, such as autonomous vehicles, can be safely and inexpensively conducted offline, i.e. by computing model prediction error over a pre-collected validation dataset with ground-truth annotations. However, extrapolating from offline model performance to online settings remains a challenge. In these settings, seemingly minor errors can compound and result in test-time infractions or collisions. This relationship is understudied, particularly across diverse closed-loop metrics and complex urban maneuvers. In this work, we revisit this undervalued question in policy evaluation through an extensive set of experiments across diverse conditions and metrics. Based on analysis in simulation, we find an even worse correlation between offline and online settings than reported by prior studies, casting doubts on the validity of current evaluation practices and metrics for driving policies. Next, we bridge the gap between offline and online evaluation. We investigate an offline metric based on epistemic uncertainty, which aims to capture events that are likely to cause errors in closed-loop settings. The resulting metric achieves over 13% improvement in correlation compared to previous offline metrics. We further validate the generalization of our findings beyond the simulation environment in real-world settings, where even greater gains are observed.

📄 PDF Abstract BibTeX arXiv:2510.08571

Code (0)

등록된 구현이 없습니다.

Tasks

Autonomous VehiclesAutonomous Driving

Similar Papers 제목 키워드 기반

On Offline Evaluation of Vision-based Driving Models

2018-09-13 · ECCV 2018 9 · Felipe Codevilla, Antonio M. López, Vladlen Koltun, Alexey Dosovitskiy

Autonomous driving models should ideally be evaluated by deploying them on a fleet of physical vehicles in the real world. Unfortunately, this approach is not practical for the vast majority of researchers. An attractive…

Autonomous Driving

On Offline Evaluation of 3D Object Detection for Autonomous Driving

2023-08-24 · Tim Schreier, Katrin Renz, Andreas Geiger, Kashyap Chitta

Prior work in 3D object detection evaluates models using offline metrics like average precision since closed-loop online evaluation on the downstream driving task is costly. However, it is unclear how indicative offline …

3D Object DetectionAutonomous DrivingObjectobject-detection+1

Real-World Perturbation Testing of Autonomous Driving Systems

2026-07-06 · Stefano Carlo Lambertenghi, Matthias Weil, Andrea Stocco arxiv

Autonomous Driving Systems (ADS) must operate reliably under diverse conditions, yet representative data for rare or adverse scenarios is difficult to obtain. Perturbation-based testing is widely used to assess robustnes…

Autonomous Driving

Driving with A Thousand Faces: A Benchmark for Closed-Loop Personalized End-to-End Autonomous Driving

2026-02-21 · Xiaoru Dong, Ruiqin Li, Xiao Han, Zhenxuan Wu 외 arxiv

Human driving behavior is inherently diverse, yet most end-to-end autonomous driving (E2E-AD) systems learn a single average driving style, neglecting individual differences. Achieving personalized E2E-AD faces challenge…

Autonomous Driving

Large Scale Autonomous Driving Scenarios Clustering with Self-supervised Feature Extraction

2021-03-30 · Jinxin Zhao, Jin Fang, Zhixian Ye, Liangjun Zhang

The clustering of autonomous driving scenario data can substantially benefit the autonomous driving validation and simulation systems by improving the simulation tests' completeness and fidelity. This article proposes a …

Autonomous DrivingClusteringData AugmentationFeature Compression