paper-with-me

Papers

PredictaBoard: Benchmarking LLM Score Predictability

2025-02-20 · Lorenzo Pacchiardi, Konstantinos Voudouris, Ben Slater, Fernando Martínez-Plumed, José Hernández-Orallo, Lexin Zhou, Wout Schellaert

Despite possessing impressive skills, Large Language Models (LLMs) often fail unpredictably, demonstrating inconsistent success in even basic common sense reasoning tasks. This unpredictability poses a significant challenge to ensuring their safe deployment, as identifying and operating within a reliable "safe zone" is essential for mitigating risks. To address this, we present PredictaBoard, a novel collaborative benchmarking framework designed to evaluate the ability of score predictors (referred to as assessors) to anticipate LLM errors on specific task instances (i.e., prompts) from existing datasets. PredictaBoard evaluates pairs of LLMs and assessors by considering the rejection rate at different tolerance errors. As such, PredictaBoard stimulates research into developing better assessors and making LLMs more predictable, not only with a higher average performance. We conduct illustrative experiments using baseline assessors and state-of-the-art LLMs. PredictaBoard highlights the critical need to evaluate predictability alongside performance, paving the way for safer AI systems where errors are not only minimised but also anticipated and effectively mitigated. Code for our benchmark can be found at https://github.com/Kinds-of-Intelligence-CFI/PredictaBoard

📄 PDF Abstract BibTeX arXiv:2502.14445

Code (1)

Kinds-of-Intelligence-CFI/PredictaBoard 공식 구현

Tasks

BenchmarkingCommon Sense Reasoning

Similar Papers 제목 키워드 기반

The Benchmarking Epistemology: Construct Validity for Evaluating Machine Learning Models

2025-10-27 · Timo Freiesleben, Sebastian Zezulka arxiv

Predictive benchmarking, the evaluation of machine learning models based on predictive performance and competitive ranking, is a central epistemic practice in machine learning research and an increasingly prominent metho…

Image Classification

Measuring the Predictability of Recommender Systems using Structural Complexity Metrics

2024-04-12 · Alfonso Valderrama, Andrés Abeliuk

Recommender systems (RS) are central to the filtering and curation of online content. These algorithms predict user ratings for unseen items based on past preferences. Despite their importance, the innate predictability …

Recommendation Systems

Beyond Model Ranking: Predictability-Aligned Evaluation for Time Series Forecasting

2025-09-27 · Wanjin Feng, Yuan Yuan, Jingtao Ding, Yong Li arxiv

In the era of increasingly complex AI models for time series forecasting, progress is often measured by marginal improvements on benchmark leaderboards. However, this approach suffers from a fundamental flaw: standard ev…

Time Series Forecasting

OpenTraj: Assessing Prediction Complexity in Human Trajectories Datasets

2020-10-02 · Javad Amirian, Bingqing Zhang, Francisco Valente Castro, Juan Jose Baldelomar 외

Human Trajectory Prediction (HTP) has gained much momentum in the last years and many solutions have been proposed to solve it. Proper benchmarking being a key issue for comparing methods, this paper addresses the questi…

BenchmarkingPredictionSelf-Driving CarsTrajectory Forecasting+1

Hierarchy of extreme-event predictability in turbulence revealed by machine learning

2026-03-14 · Yuxuan Yang, Chenyu Dong, Gianmarco Mengaldo arxiv

Extreme-event predictability in turbulence is strongly state dependent, yet event-by-event predictability horizons are difficult to quantify without access to governing equations or costly perturbation ensembles. Here we…