paper-with-me

Papers

Confident Rankings with Fewer Items: Adaptive LLM Evaluation with Continuous Scores

2026-01-20 · Esma Balkır, Alice Pernthaller, Marco Basaldella, José Hernández-Orallo, Nigel Collier arxiv

Computerized Adaptive Testing (CAT) has proven effective for efficient LLM evaluation on multiple-choice benchmarks, but modern LLM evaluation increasingly relies on generation tasks where outputs are scored continuously rather than marked correct/incorrect. We present a principled extension of IRT-based adaptive testing to continuous bounded scores (ROUGE, BLEU, LLM-as-a-Judge) by replacing the Bernoulli response distribution with a heteroskedastic normal distribution. Building on this, we introduce an uncertainty aware ranker with adaptive stopping criteria that achieves reliable model ranking while testing as few items and as cheaply as possible. We validate our method on five benchmarks spanning n-gram-based, embedding-based, and LLM-as-judge metrics. Our method uses 2% of the items while improving ranking correlation by 0.12 τ over random sampling, with 95% accuracy on confident predictions.

📄 PDF Abstract BibTeX arXiv:2601.13885

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Learning when to rank: Estimation of partial rankings from sparse, noisy comparisons

2025-01-05 · Sebastian Morel-Balbi, Alec Kirkley

A common task arising in various domains is that of ranking items based on the outcomes of pairwise comparisons, from ranking players and teams in sports to ranking products or brands in marketing studies and recommendat…

MarketingRecommendation Systems

Consensus measure of rankings

2017-04-27 · Zhiwei Lin, Yi Li, Xiaolian Guo

A ranking is an ordered sequence of items, in which an item with higher ranking score is more preferred than the items with lower ranking scores. In many information systems, rankings are widely used to represent the pre…

The few-get-richer: a surprising consequence of popularity-based rankings

2019-02-07 · Fabrizio Germano, Vicenç Gómez, Gaël Le Mens

Ranking algorithms play a crucial role in online platforms ranging from search engines to recommender systems. In this paper, we identify a surprising consequence of popularity-based rankings: the fewer the items reporti…

MisinformationRecommendation Systems

Data-Driven Relevance Judgments for Ranking Evaluation

2016-12-19 · Moniz Nuno, Torgo Luís, Vinagre João

Ranking evaluation metrics are a fundamental element of design and improvement efforts in information retrieval. We observe that most popular metrics disregard information portrayed in the scores used to derive rankings,…

Information RetrievalRetrieval

Decomposition and Interleaving for Variance Reduction of Post-click Metrics

2023-05-31 · Kojiro Iizuka, Yoshifumi Seki, Makoto P. Kato

In this study, we propose an efficient method for comparing the post-click metric (e.g., dwell time and conversion rate) of multiple rankings in online experiments. The proposed method involves (1) the decomposition of t…