paper-with-me

홈 › Papers

Does the evaluation stand up to evaluation? A first-principle approach to the evaluation of classifiers

2023-02-21 · K. Dyrland, A. S. Lundervold, P. G. L. Porta Mana

How can one meaningfully make a measurement, if the meter does not conform to any standard and its scale expands or shrinks depending on what is measured? In the present work it is argued that current evaluation practices for machine-learning classifiers are affected by this kind of problem, leading to negative consequences when classifiers are put to real use; consequences that could have been avoided. It is proposed that evaluation be grounded on Decision Theory, and the implications of such foundation are explored. The main result is that every evaluation metric must be a linear combination of confusion-matrix elements, with coefficients - "utilities" - that depend on the specific classification problem. For binary classification, the space of such possible metrics is effectively two-dimensional. It is shown that popular metrics such as precision, balanced accuracy, Matthews Correlation Coefficient, Fowlkes-Mallows index, F1-measure, and Area Under the Curve are never optimal: they always give rise to an in-principle avoidable fraction of incorrect evaluations. This fraction is even larger than would be caused by the use of a decision-theoretic metric with moderately wrong coefficients.

📄 PDF Abstract BibTeX arXiv:2302.12006

Code (0)

등록된 구현이 없습니다.

Tasks

Binary Classification

Similar Papers 제목 키워드 기반

Sharing is CAIRing: Characterizing Principles and Assessing Properties of Universal Privacy Evaluation for Synthetic Tabular Data

2023-12-19 · Tobias Hyrup, Anton Danholt Lautrup, Arthur Zimek, Peter Schneider-Kamp

Data sharing is a necessity for innovative progress in many domains, especially in healthcare. However, the ability to share data is hindered by regulations protecting the privacy of natural persons. Synthetic tabular da…

Privacy Preserving

Interactive Evaluation Requires a Design Science

2026-05-18 · Keyang Xuan, Peiyang Song, Pan Lu, Pengrui Han 외 arxiv

AI evaluation is undergoing a structural change. Large language models (LLMs) are increasingly deployed as systems that act over time through tools, environments, users, and other agents, while many evaluation practices …

Realistic Evaluation Principles for Cross-document Coreference Resolution

2021-06-08 · Joint Conference on Lexical and Computational Semantics 2021 · Arie Cattan, Alon Eirew, Gabriel Stanovsky, Mandar Joshi 외

We point out that common evaluation practices for cross-document coreference resolution have been unrealistically permissive in their assumed settings, yielding inflated results. We propose addressing this issue via two …

coreference-resolutionCoreference ResolutionCross Document Coreference Resolution

No Metric to Rule Them All: Toward Principled Evaluations of Graph-Learning Datasets

2025-02-04 · Corinna Coupette, Jeremy Wayland, Emily Simons, Bastian Rieck

Benchmark datasets have proved pivotal to the success of graph learning, and good benchmark datasets are crucial to guide the development of the field. Recent research has highlighted problems with graph-learning dataset…

AllBenchmarkingGraph Learning

Towards a Decomposable Metric for Explainable Evaluation of Text Generation from AMR

2020-08-20 · EACL 2021 2 · Juri Opitz, Anette Frank

Systems that generate natural language text from abstract meaning representations such as AMR are typically evaluated using automatic surface matching metrics that compare the generated texts to reference texts from whic…

Abstract Meaning RepresentationSentenceText Generation