Statistically Reliable LLM-Based Ranking Evaluation via Prediction-Powered Inference
With PRECISE, we extended Prediction-Powered Inference to produce bias-corrected estimates of ranking evaluation metrics by combining a small human-labeled set with a large LLM-judged set. PPI is provably unbiased regardless of the LLM judge's error profile. We make it applicable to hierarchical metrics like Precision@K, where annotations are per-document but the metric is per-query, by reducing the output-space computation from O(2^|C|) to O(2^K). On the ESCI benchmark, augmenting 30 human annotations with Claude 3 Sonnet judgments reduces the standard error of Precision@4 estimates from 4.45 to 3.50 (a 21% relative reduction). In a production system, our framework correctly identified the best of three system variants from 100 human labels and 2 hours of domain-expert annotation; A/B testing confirmed this ranking with +407 bps in daily sales.
Code (0)
등록된 구현이 없습니다.
Similar Papers 제목 키워드 기반
Which Metrics Save the Most Human Annotation? Prediction-Powered Evaluation and Meta-Evaluation
Across various non-verifiable tasks, human evaluation is reliable but expensive, while automatic metrics are more scalable but often biased. Building on prediction-powered inference (PPI), we propose prediction-powered e…
Reliable Confidence Intervals for Information Retrieval Evaluation Using Generative A.I
The traditional evaluation of information retrieval (IR) systems is generally very costly as it requires manual relevance annotation from human experts. Recent advancements in generative artificial intelligence -- specif…
Information RetrievalRetrievalRank Intervals for Leaderboards: A Hierarchical Framework for Model Evaluation
Pretrained models are often evaluated on multi-task leaderboards to measure their applicability in diverse contexts. However, current methods for aggregating performance across tasks into leaderboard-level rankings do no…
PRECISE: Reducing the Bias of LLM Evaluations Using Prediction-Powered Ranking Estimation
Evaluating the quality of search, ranking and RAG systems traditionally requires a significant number of human relevance annotations. In recent times, several deployed systems have explored the usage of Large Language Mo…
Federated Prediction-Powered Inference from Decentralized Data
In various domains, the increasing application of machine learning allows researchers to access inexpensive predictive data, which can be utilized as auxiliary data for statistical inference. Although such data are often…
Federated LearningPredictionvalid