paper-with-me

Papers

A Judge-Aware Ranking Framework for Evaluating Large Language Models without Ground Truth

2026-01-29 · Mingyuan Xu, Xinzi Tan, Jiawei Wu, Doudou Zhou arxiv

Evaluating large language models (LLMs) on open-ended tasks without ground-truth labels is increasingly done via the LLM-as-a-judge paradigm. A critical but under-modeled issue is that judge LLMs differ substantially in reliability; treating all judges equally can yield biased leaderboards and misleading uncertainty estimates. More data can make evaluation more confidently wrong under misspecified aggregation. We propose a judge-aware ranking framework that extends the Bradley-Terry-Luce model by introducing judge-specific discrimination parameters, jointly estimating latent model quality and judge reliability from pairwise comparisons without reference labels. We establish identifiability up to natural normalizations and prove consistency and asymptotic normality of the maximum likelihood estimator, enabling confidence intervals for score differences and rank comparisons. Across multiple public benchmarks and a newly collected dataset, our method improves agreement with human preferences, achieves higher data efficiency than unweighted baselines, and produces calibrated uncertainty quantification for LLM rankings.

📄 PDF Abstract BibTeX arXiv:2601.21817

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

AutoJudger: An Agent-Driven Framework for Efficient Benchmarking of MLLMs

2025-05-27 · Xuanwen Ding, Chengjun Pan, Zejun Li, Jiwen Zhang 외

Evaluating multimodal large language models (MLLMs) is increasingly expensive, as the growing size and cross-modality complexity of benchmarks demand significant scoring efforts. To tackle with this difficulty, we introd…

BenchmarkingQuestion Selection

Meaning in Order, Order in Meaning: Semantic R-precision for Keyphrase Evaluation

2026-06-05 · Shamira Venturini, Steffen Kinkel arxiv

Evaluating the quality of automatically generated keyphrases remains a complex challenge. Traditional metrics either rely on exact lexical matching or consider semantic similarity while ignoring prediction ranking, both …

Information RetrievalSemantic Similarity

Ask the Right Comparison:Bias-Aware Bayesian Active Top-$k$ Ranking with LLM Judges

2026-07-02 · Jian Xu, Delu Zeng, John Paisley, Qibin Zhao arxiv

Large language models (LLMs) are increasingly used as cheap, scalable judges that compare candidate outputs pairwise -- to rank responses, select models, or triage papers. Yet LLM judges are both noisy and systematically…

Bayesian Inference

Correct Looks Better: Pairwise Comparisons Reveal Accuracy Rankings

2026-06-08 · Mina Remeli, Moritz Hardt arxiv

Pairwise comparisons combined with aggregation methods like Elo have become central to evaluating generative models, yet concerns remain that they reward superficial stylistic cues or display judge biases. In a more posi…

Multi-Dimensional Evaluation of Sustainable City Trips with LLM-as-a-Judge and Human-in-the-Loop

2026-04-27 · Ashmi Banerjee, Adithi Satish, Wolfgang Wörndl, Yashar Deldjoo arxiv

Evaluating nuanced conversational travel recommendations is challenging when human annotations are costly and standard metrics ignore stakeholder-centric goals. We study LLMs-as-Judges for sustainable city-trip lists acr…