paper-with-me

Papers

Adding Error Bars to Evals: A Statistical Approach to Language Model Evaluations

2024-11-01 · Evan Miller

Evaluations are critical for understanding the capabilities of large language models (LLMs). Fundamentally, evaluations are experiments; but the literature on evaluations has largely ignored the literature from other sciences on experiment analysis and planning. This article shows researchers with some training in statistics how to think about and analyze data from language model evaluations. Conceptualizing evaluation questions as having been drawn from an unseen super-population, we present formulas for analyzing evaluation data, measuring differences between two models, and planning an evaluation experiment. We make a number of specific recommendations for running language model evaluations and reporting experiment results in a way that minimizes statistical noise and maximizes informativeness.

📄 PDF Abstract BibTeX arXiv:2411.00640

Code (0)

등록된 구현이 없습니다.

Tasks

InformativenessLanguage ModelingLanguage Modelling

Similar Papers 제목 키워드 기반

Cross-validation failure: small sample sizes lead to large error bars

2017-06-23 · Gaël Varoquaux

Predictive models ground many state-of-the-art developments in statistical brain image analysis: decoding, MVPA, searchlight, or extraction of biomarkers. The principled approach to establish their validity and usefulnes…

EvalSense: A Framework for Domain-Specific LLM (Meta-)Evaluation

2026-02-21 · Adam Dejl, Jonathan Pearson arxiv

Robust and comprehensive evaluation of large language models (LLMs) is essential for identifying effective LLM system configurations and mitigating risks associated with deploying LLMs in sensitive domains. However, trad…

Uncertainty Quantification of Surrogate Models using Conformal Prediction

2024-08-19 · Vignesh Gopakumar, Ander Gray, Joel Oskarsson, Lorenzo Zanisi 외

Data-driven surrogate models have shown immense potential as quick, inexpensive approximations to complex numerical and experimental modelling tasks. However, most surrogate models of physical systems do not quantify the…

Conformal PredictionPredictionUncertainty Quantificationvalid+1

Measuring all the noises of LLM Evals

2025-12-24 · Sida Wang arxiv

Separating signal from noise is central to experiments. Applying well-established statistical methods effectively to LLM evals requires consideration of their unique noise characteristics. We clearly define and measure t…

Xspect, estimation of the angular power spectrum by computing cross-power spectra with analytical error bars

2004-05-28 · M. Tristram, J. F. Macias-Perez, C. Renault, D. Santos

We present Xspect, a method to obtain estimates of the angular power spectrum of the Cosmic Microwave Background (CMB) temperature anisotropies including analytical error bars developed for the Archeops experiment. Cross…