paper-with-me

홈 › Papers

Measuring What Matters: Scenario-Driven Evaluation for Trajectory Predictors in Autonomous Driving

2025-12-13 · Longchao Da, David Isele, Hua Wei, Manish Saroya arxiv

Being able to anticipate the motion of surrounding agents is essential for the safe operation of autonomous driving systems in dynamic situations. While various methods have been proposed for trajectory prediction, the current evaluation practices still rely on error-based metrics (e.g., ADE, FDE), which reveal the accuracy from a post-hoc view but ignore the actual effect the predictor brings to the self-driving vehicles (SDVs), especially in complex interactive scenarios: a high-quality predictor not only chases accuracy, but should also captures all possible directions a neighbor agent might move, to support the SDVs' cautious decision-making. Given that the existing metrics hardly account for this standard, in our work, we propose a comprehensive pipeline that adaptively evaluates the predictor's performance by two dimensions: accuracy and diversity. Based on the criticality of the driving scenario, these two dimensions are dynamically combined and result in a final score for the predictor's performance. Extensive experiments on a closed-loop benchmark using real-world datasets show that our pipeline yields a more reasonable evaluation than traditional metrics by better reflecting the correlation of the predictors' evaluation with the autonomous vehicles' driving performance. This evaluation pipeline shows a robust way to select a predictor that potentially contributes most to the SDV's driving performance.

📄 PDF Abstract BibTeX arXiv:2512.12211

Code (0)

등록된 구현이 없습니다.

Tasks

Trajectory PredictionAutonomous VehiclesAutonomous Driving

Similar Papers 제목 키워드 기반

Measuring what Matters: Construct Validity in Large Language Model Benchmarks

2025-11-03 · Andrew M. Bean, Ryan Othniel Kearns, Angelika Romanou, Franziska Sofia Hafner 외 arxiv

Evaluating large language models (LLMs) is crucial for both assessing their capabilities and identifying safety or robustness issues prior to deployment. Reliably measuring abstract and complex phenomena such as 'safety'…

A Fuzzy Approach to Project Success: Measuring What Matters

2025-07-16 · João Granja-Correia, Remedios Hernández-Linares, Luca Ferranti, Arménio Rego

This paper introduces a novel approach to project success evaluation by integrating fuzzy logic into an existing construct. Traditional Likert-scale measures often overlook the context-dependent and multifaceted nature o…

Measuring What AI Systems Might Do: Towards A Measurement Science in AI

2026-02-10 · Konstantinos Voudouris, Mirko Thalmann, Alex Kipnis, José Hernández-Orallo 외 arxiv

Scientists, policy-makers, business leaders, and members of the public care about what modern artificial intelligence systems are disposed to do. Yet terms such as capabilities, propensities, skills, values, and abilitie…

Mask What Matters: Controllable Text-Guided Masking for Self-Supervised Medical Image Analysis

2025-09-27 · Ruilang Wang, Shuotong Xu, Bowen Liu, Runlin Huang 외 arxiv

The scarcity of annotated data in specialized domains such as medical imaging presents significant challenges to training robust vision models. While self-supervised masked image modeling (MIM) offers a promising solutio…

Self-Supervised LearningRepresentation Learning

Measuring What Matters: Connecting AI Ethics Evaluations to System Attributes, Hazards, and Harms

2025-10-11 · Shalaleh Rismani, Renee Shelby, Leah Davis, Negar Rostamzadeh 외 arxiv

Over the past decade, an ecosystem of measures has emerged to evaluate the social and ethical implications of AI systems, largely shaped by high-level ethics principles. These measures are developed and used in fragmente…