paper-with-me

홈 › Papers

Statistically Reliable LLM-Based Ranking Evaluation via Prediction-Powered Inference

2026-06-03 · Abhishek Divekar arxiv

With PRECISE, we extended Prediction-Powered Inference to produce bias-corrected estimates of ranking evaluation metrics by combining a small human-labeled set with a large LLM-judged set. PPI is provably unbiased regardless of the LLM judge's error profile. We make it applicable to hierarchical metrics like Precision@K, where annotations are per-document but the metric is per-query, by reducing the output-space computation from O(2^|C|) to O(2^K). On the ESCI benchmark, augmenting 30 human annotations with Claude 3 Sonnet judgments reduces the standard error of Precision@4 estimates from 4.45 to 3.50 (a 21% relative reduction). In a production system, our framework correctly identified the best of three system variants from 100 human labels and 2 hours of domain-expert annotation; A/B testing confirmed this ranking with +407 bps in daily sales.

📄 PDF Abstract BibTeX arXiv:2606.05308

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Which Metrics Save the Most Human Annotation? Prediction-Powered Evaluation and Meta-Evaluation

2026-08-27 · Mingqi Gao, Anthony Sicilia, Weiyan Shi arxiv

Across various non-verifiable tasks, human evaluation is reliable but expensive, while automatic metrics are more scalable but often biased. Building on prediction-powered inference (PPI), we propose prediction-powered e…

Reliable Confidence Intervals for Information Retrieval Evaluation Using Generative A.I

2024-07-02 · Harrie Oosterhuis, Rolf Jagerman, Zhen Qin, Xuanhui Wang 외

The traditional evaluation of information retrieval (IR) systems is generally very costly as it requires manual relevance annotation from human experts. Recent advancements in generative artificial intelligence -- specif…

Information RetrievalRetrieval

Rank Intervals for Leaderboards: A Hierarchical Framework for Model Evaluation

2026-06-07 · Bitya Neuhof, Yuval Benjamini arxiv

Pretrained models are often evaluated on multi-task leaderboards to measure their applicability in diverse contexts. However, current methods for aggregating performance across tasks into leaderboard-level rankings do no…

PRECISE: Reducing the Bias of LLM Evaluations Using Prediction-Powered Ranking Estimation

2026-01-26 · Abhishek Divekar, Anirban Majumder arxiv

Evaluating the quality of search, ranking and RAG systems traditionally requires a significant number of human relevance annotations. In recent times, several deployed systems have explored the usage of Large Language Mo…

Federated Prediction-Powered Inference from Decentralized Data

2024-09-03 · Ping Luo, Xiaoge Deng, Ziqing Wen, Tao Sun 외

In various domains, the increasing application of machine learning allows researchers to access inexpensive predictive data, which can be utilized as auxiliary data for statistical inference. Although such data are often…

Federated LearningPredictionvalid