paper-with-me

홈 › Papers

Accounting for multiplicity in machine learning benchmark performance

2023-03-10 · Kajsa Møllersen, Einar Holsbø

Machine learning methods are commonly evaluated and compared by their performance on data sets from public repositories. This allows for multiple methods, oftentimes several thousands, to be evaluated under identical conditions and across time. The highest ranked performance on a problem is referred to as state-of-the-art (SOTA) performance, and is used, among other things, as a reference point for publication of new methods. Using the highest-ranked performance as an estimate for SOTA is a biased estimator, giving overly optimistic results. The mechanisms at play are those of multiplicity, a topic that is well-studied in the context of multiple comparisons and multiple testing, but has, as far as the authors are aware of, been nearly absent from the discussion regarding SOTA estimates. The optimistic state-of-the-art estimate is used as a standard for evaluating new methods, and methods with substantial inferior results are easily overlooked. In this article, we provide a probability distribution for the case of multiple classifiers so that known analyses methods can be engaged and a better SOTA estimate can be provided. We demonstrate the impact of multiplicity through a simulated example with independent classifiers. We show how classifier dependency impacts the variance, but also that the impact is limited when the accuracy is high. Finally, we discuss three real-world examples; Kaggle competitions that demonstrate various aspects.

📄 PDF Abstract BibTeX arXiv:2303.07272

Code (1)

3inar/ninety-nine 공식 구현

Methods 이 논문이 사용한 방법론

AWARE We propose to theoretically and empirically examine the effect of incorporating weighting schemes into walk-aggregating GNNs. To this end, we propose a simple, interpretable, and…

Similar Papers 제목 키워드 기반

An Empirical Investigation into Benchmarking Model Multiplicity for Trustworthy Machine Learning: A Case Study on Image Classification

2023-11-24 · Prakhar Ganesh

Deep learning models have proven to be highly successful. Yet, their over-parameterization gives rise to model multiplicity, a phenomenon in which multiple models achieve similar performance but exhibit distinct underlyi…

Benchmarkingimage-classificationImage ClassificationModel Selection

Decomposing Observational Multiplicity in Decision Trees: Leaf and Structural Regret

2026-03-12 · Mustafa Cavus arxiv

Many machine learning tasks admit multiple models that perform almost equally well, a phenomenon known as predictive multiplicity. A fundamental source of this multiplicity is observational multiplicity, which arises fro…

On Arbitrary Predictions from Equally Valid Models

2025-07-25 · Sarah Lockfisch, Kristian Schwethelm, Martin Menten, Rickmer Braren 외 arxiv

Model multiplicity refers to the existence of multiple machine learning models that describe the data equally well but may produce different predictions on individual samples. In medicine, these models can admit conflict…

The Role of Hyperparameters in Predictive Multiplicity

2025-03-13 · Mustafa Cavus, Katarzyna Woźnica, Przemysław Biecek

This paper investigates the critical role of hyperparameters in predictive multiplicity, where different machine learning models trained on the same dataset yield divergent predictions for identical inputs. These inconsi…

FairnessHyperparameter OptimizationPrediction

Accounting for Model Uncertainty in Algorithmic Discrimination

2021-05-10 · Junaid Ali, Preethi Lahoti, Krishna P. Gummadi

Traditional approaches to ensure group fairness in algorithmic decision making aim to equalize ``total'' error rates for different subgroups in the population. In contrast, we argue that the fairness approaches should in…

Decision MakingFairnessmodel