Adding Error Bars to Evals: A Statistical Approach to Language Model Evaluations
Evaluations are critical for understanding the capabilities of large language models (LLMs). Fundamentally, evaluations are experiments; but the literature on evaluations has largely ignored the literature from other sciences on experiment analysis and planning. This article shows researchers with some training in statistics how to think about and analyze data from language model evaluations. Conceptualizing evaluation questions as having been drawn from an unseen super-population, we present formulas for analyzing evaluation data, measuring differences between two models, and planning an evaluation experiment. We make a number of specific recommendations for running language model evaluations and reporting experiment results in a way that minimizes statistical noise and maximizes informativeness.
Code (0)
등록된 구현이 없습니다.
Tasks
InformativenessLanguage ModelingLanguage ModellingSimilar Papers 제목 키워드 기반
Cross-validation failure: small sample sizes lead to large error bars
Predictive models ground many state-of-the-art developments in statistical brain image analysis: decoding, MVPA, searchlight, or extraction of biomarkers. The principled approach to establish their validity and usefulnes…
EvalSense: A Framework for Domain-Specific LLM (Meta-)Evaluation
Robust and comprehensive evaluation of large language models (LLMs) is essential for identifying effective LLM system configurations and mitigating risks associated with deploying LLMs in sensitive domains. However, trad…
Uncertainty Quantification of Surrogate Models using Conformal Prediction
Data-driven surrogate models have shown immense potential as quick, inexpensive approximations to complex numerical and experimental modelling tasks. However, most surrogate models of physical systems do not quantify the…
Conformal PredictionPredictionUncertainty Quantificationvalid+1Measuring all the noises of LLM Evals
Separating signal from noise is central to experiments. Applying well-established statistical methods effectively to LLM evals requires consideration of their unique noise characteristics. We clearly define and measure t…
Xspect, estimation of the angular power spectrum by computing cross-power spectra with analytical error bars
We present Xspect, a method to obtain estimates of the angular power spectrum of the Cosmic Microwave Background (CMB) temperature anisotropies including analytical error bars developed for the Archeops experiment. Cross…