Statistical Multicriteria Evaluation of LLM-Generated Text
Assessing the quality of LLM-generated text remains a fundamental challenge in natural language processing. Current evaluation approaches often rely on isolated metrics or simplistic aggregations that fail to capture the nuanced trade-offs between coherence, diversity, fluency, and other relevant indicators of text quality. In this work, we adapt a recently proposed framework for statistical inference based on Generalized Stochastic Dominance (GSD) that addresses three critical limitations in existing benchmarking methodologies: the inadequacy of single-metric evaluation, the incompatibility between cardinal automatic metrics and ordinal human judgments, and the lack of inferential statistical guarantees. The GSD-front approach enables simultaneous evaluation across multiple quality dimensions while respecting their different measurement scales, building upon partial orders of decoding strategies, thus avoiding arbitrary weighting of the involved metrics. By applying this framework to evaluate common decoding strategies against human-generated text, we demonstrate its ability to identify statistically significant performance differences while accounting for potential deviations from the i.i.d. assumption of the sampling design.
Code (1)
Tasks
BenchmarkingDiversitySimilar Papers 제목 키워드 기반
Application of independent component analysis and TOPSIS to deal with dependent criteria in multicriteria decision problems
A vast number of multicriteria decision making methods have been developed to deal with the problem of ranking a set of alternatives evaluated in a multicriteria fashion. Very often, these methods assume that the evaluat…
blind source separationDecision MakingTowards Better Open-Ended Text Generation: A Multicriteria Evaluation Framework
Open-ended text generation has become a prominent task in natural language processing due to the rise of powerful (large) language models. However, evaluating the quality of these models and the employed decoding strateg…
BenchmarkingDiversityModel SelectionText GenerationA fairness-aware extension of Stochastic Multicriteria Acceptability Analysis for ranking
Fairness has become a central concern in ranking problems involving individuals or social groups, particularly under the Responsible Artificial Intelligence agenda. In Multi-Criteria Decision Analysis, Stochastic Multicr…
Multicriteria Analysis Model in Sustainable Corn Farming Area Planning
This study aims to develop a framework for multicriteria analysis to evaluate alternatives for sustainable corn agricultural area planning, considering the integration of ecological, economic, and social aspects as pilla…
Data IntegrationA multiscale and multicriteria Generative Adversarial Network to synthesize 1-dimensional turbulent fields
This article introduces a new Neural Network stochastic model to generate a 1-dimensional stochastic field with turbulent velocity statistics. Both the model architecture and training procedure ground on the Kolmogorov a…
Generative Adversarial Network