paper-with-me

홈 › Papers

Rank Intervals for Leaderboards: A Hierarchical Framework for Model Evaluation

2026-06-07 · Bitya Neuhof, Yuval Benjamini arxiv

Pretrained models are often evaluated on multi-task leaderboards to measure their applicability in diverse contexts. However, current methods for aggregating performance across tasks into leaderboard-level rankings do not address the uncertainty and variability at the task level. While recent works have proposed interval-based model rankings, the principled aggregation of uncertainty from individual tasks to leaderboard-level rankings remains unaddressed, and variation in models' performance across tasks is frequently obscured. In this work, we introduce a hierarchical framework that constructs model rank intervals with statistical guarantees at both levels: task-level rank confidence intervals from pairwise comparisons, and leaderboard-level rank prediction intervals using a conformal approach. This enables reliable quantification of model rank for each observed task and for new potential tasks. Experiments on simulated data and the TabArena and PromptEval (MMLU) benchmarks show that our method yields statistically valid and informative intervals, enabling reliable, uncertainty-aware model ranking on leaderboards.

📄 PDF Abstract BibTeX arXiv:2606.08679

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

A Unified Perturbation Framework for Analyzing Leaderboard Stability and Manipulation

2026-05-15 · Hosna Oyarhoseini, Jimmy Lin, Amir-Hossein Karimi arxiv

Evaluation leaderboards such as LMArena play a central role in benchmarking large language models by aggregating pairwise human preferences into model rankings, yet the robustness of these rankings remains poorly underst…

Evaluation Examples are not Equally Informative: How should that change NLP Leaderboards?

2021-08-01 · ACL 2021 5 · Pedro Rodriguez, Joe Barrow, Alexander Miserlis Hoyle, John P. Lalor 외

Leaderboards are widely used in NLP and push the field forward. While leaderboards are a straightforward ranking of NLP models, this simplicity can mask nuances in evaluation items (examples) and subjects (NLP models). R…

A Judge-Aware Ranking Framework for Evaluating Large Language Models without Ground Truth

2026-01-29 · Mingyuan Xu, Xinzi Tan, Jiawei Wu, Doudou Zhou arxiv

Evaluating large language models (LLMs) on open-ended tasks without ground-truth labels is increasingly done via the LLM-as-a-judge paradigm. A critical but under-modeled issue is that judge LLMs differ substantially in …

Who Defines "Best"? Towards Interactive, User-Defined Evaluation of LLM Leaderboards

2026-04-23 · Minji Jung, Minjae Lee, Yejin Kim, Sarang Choi 외 arxiv

LLM leaderboards are widely used to compare models and guide deployment decisions. However, leaderboard rankings are shaped by evaluation priorities set by benchmark designers, rather than by the diverse goals and constr…

Prompt-Dependent Ranking of Large Language Models with Uncertainty Quantification

2026-02-11 · Angel Rodrigo Avelar Menendez, Yufeng Liu, Xiaowu Dai arxiv

Rankings derived from pairwise comparisons are central to many economic and computational systems. In the context of large language models (LLMs), rankings are typically constructed from human preference data and present…