Statistical Multicriteria Benchmarking via the GSD-Front
Given the vast number of classifiers that have been (and continue to be) proposed, reliable methods for comparing them are becoming increasingly important. The desire for reliability is broken down into three main aspects: (1) Comparisons should allow for different quality metrics simultaneously. (2) Comparisons should take into account the statistical uncertainty induced by the choice of benchmark suite. (3) The robustness of the comparisons under small deviations in the underlying assumptions should be verifiable. To address (1), we propose to compare classifiers using a generalized stochastic dominance ordering (GSD) and present the GSD-front as an information-efficient alternative to the classical Pareto-front. For (2), we propose a consistent statistical estimator for the GSD-front and construct a statistical test for whether a (potentially new) classifier lies in the GSD-front of a set of state-of-the-art classifiers. For (3), we relax our proposed test using techniques from robust statistics and imprecise probabilities. We illustrate our concepts on the benchmark suite PMLB and on the platform OpenML.
Code (0)
등록된 구현이 없습니다.
Tasks
BenchmarkingMethods 이 논문이 사용한 방법론
Similar Papers 제목 키워드 기반
Statistical Multicriteria Evaluation of LLM-Generated Text
Assessing the quality of LLM-generated text remains a fundamental challenge in natural language processing. Current evaluation approaches often rely on isolated metrics or simplistic aggregations that fail to capture the…
BenchmarkingDiversityApplication of independent component analysis and TOPSIS to deal with dependent criteria in multicriteria decision problems
A vast number of multicriteria decision making methods have been developed to deal with the problem of ranking a set of alternatives evaluated in a multicriteria fashion. Very often, these methods assume that the evaluat…
blind source separationDecision MakingRobust measurement of innovation performances in Europe with a hierarchy of interacting composite indicators
For long time the measurement of innovation has been in the forefront of policy makers' and researchers' agenda worldwide. Therefore, there is an ongoing debate about which indicators should be used to measure innovation…
BenchmarkingDecision MakingComparative Statics in Multicriteria Search Models
McCall (1970) examines the search behaviour of an infinitely-lived and risk-neutral job seeker maximizing her lifetime earnings by accepting or rejecting real-valued scalar wage offers. In practice, job offers have multi…
Towards Better Open-Ended Text Generation: A Multicriteria Evaluation Framework
Open-ended text generation has become a prominent task in natural language processing due to the rise of powerful (large) language models. However, evaluating the quality of these models and the employed decoding strateg…
BenchmarkingDiversityModel SelectionText Generation