paper-with-me

홈 › Papers

PRECISE: Reducing the Bias of LLM Evaluations Using Prediction-Powered Ranking Estimation

2026-01-26 · Abhishek Divekar, Anirban Majumder arxiv

Evaluating the quality of search, ranking and RAG systems traditionally requires a significant number of human relevance annotations. In recent times, several deployed systems have explored the usage of Large Language Models (LLMs) as automated judges for this task while their inherent biases prevent direct use for metric estimation. We present a statistical framework extending Prediction-Powered Inference (PPI) that combines minimal human annotations with LLM judgments to produce reliable estimates of metrics which require sub-instance annotations. Our method requires as few as 100 human-annotated queries and 10,000 unlabeled examples, reducing annotation requirements significantly compared to traditional approaches. We formulate our proposed framework (PRECISE) for inference of relevance uplift for an LLM-based query reformulation application, extending PPI to sub-instance annotations at the query-document level. By reformulating the metric-integration space, we reduced the computational complexity from O(2^|C|) to O(2^K), where |C| represents corpus size (in order of millions). Detailed experiments across prominent retrieval datasets demonstrate that our method reduces the variance of estimates for the business-critical Precision@K metric, while effectively correcting for LLM bias in low-resource settings.

📄 PDF Abstract BibTeX arXiv:2601.18777

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Statistically Reliable LLM-Based Ranking Evaluation via Prediction-Powered Inference

2026-06-03 · Abhishek Divekar arxiv

With PRECISE, we extended Prediction-Powered Inference to produce bias-corrected estimates of ranking evaluation metrics by combining a small human-labeled set with a large LLM-judged set. PPI is provably unbiased regard…

Which Metrics Save the Most Human Annotation? Prediction-Powered Evaluation and Meta-Evaluation

2026-08-27 · Mingqi Gao, Anthony Sicilia, Weiyan Shi arxiv

Across various non-verifiable tasks, human evaluation is reliable but expensive, while automatic metrics are more scalable but often biased. Building on prediction-powered inference (PPI), we propose prediction-powered e…

Prediction-Powered Active Testing

2026-07-09 · Kianoosh Ashouritaklimi, Valentin Kilian, Daolang Huang, Tom Rainforth 외 arxiv

Active testing provides a label--efficient approach to risk estimation by adaptively selecting which test points should be labelled. However, existing estimators fail to exploit the informative predictions of powerful bl…

Adaptive Prediction-Powered AutoEval with Reliability and Efficiency Guarantees

2025-05-24 · Sangwoo Park, Matteo Zecchin, Osvaldo Simeone

Selecting artificial intelligence (AI) models, such as large language models (LLMs), from multiple candidates requires accurate performance estimation. This is ideally achieved through empirical evaluations involving abu…

Quantization

Empirical Bayes Rebiasing

2026-05-08 · Wanyi Ling, Sida Li, Junming Guan, Nikolaos Ignatiadis arxiv

We study methods for simultaneous analysis of many noisy and biased estimates, each paired with an even noisier estimate of its own bias. The analyst's goal is to construct short calibrated intervals for each parameter. …