paper-with-me

홈 › Papers

Efficient Inference for Noisy LLM-as-a-Judge Evaluation

2026-01-08 · Yiqun T Chen, Sizhu Lu, Sijia Li, Moran Guo, Shengyi Li arxiv

Large language models (LLMs) are increasingly used as automatic evaluators of generative AI outputs, a paradigm often referred to as "LLM-as-a-judge." In practice, LLM judges are imperfect predictions for the underlying truth and can exhibit systematic, non-random errors. Two main approaches have recently been proposed to address this issue: (i) direct measurementerror correction based on misclassification models such as Rogan-Gladen-style estimators, and (ii) surrogate-outcome approaches such as prediction-powered inference (PPI), which correct bias by calibrating prediction residuals on a small set of gold-standard human labels. In this paper, we systematically study the performance of these two approaches for estimating mean parameters (e.g., average benchmark scores or pairwise win rates). Leveraging tools from semiparametric efficiency theory, we unify the two classes of estimators by deriving explicit forms of efficient influence function (EIF)-based efficient estimators and characterize conditions under which PPI-style estimators attain strictly smaller asymptotic variance than measurement-error corrections. We verify our theoretical results in simulations and demonstrate the methods on real-data examples. We provide an implementation of the benchmarked methods and comparison utilities at https://github.com/yiqunchen/debias-llm-as-a-judge.

📄 PDF Abstract BibTeX arXiv:2601.05420

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Noisy but Valid: Robust Statistical Evaluation of LLMs with Imperfect Judges

2026-01-28 · Chen Feng, Minghe Shen, Ananth Balashankar, Carsten Gerner-Beuerle 외 arxiv

Reliable certification of Large Language Models (LLMs)-verifying that failure rates are below a safety threshold-is critical yet challenging. While "LLM-as-a-Judge" offers scalability, judge imperfections, noise, and bia…

Distribution-Calibrated Inference Time Compute for Thinking LLM-as-a-Judge

2025-12-02 · Hamid Dadkhahi, Firas Trabelsi, Parker Riley, Juraj Juraska 외 arxiv

Thinking Large Language Models (LLMs) used as judges for pairwise preferences remain noisy at the single-sample level, and common aggregation rules (majority vote, soft self-consistency, or instruction-based self-aggrega…

JAF: Judge Agent Forest

2026-01-29 · Sahil Garg, Brad Cheezum, Sridhar Dutta, Vishal Agarwal arxiv

Judge agents are fundamental to agentic AI frameworks: they provide automated evaluation, and enable iterative self-refinement of reasoning processes. We introduce JAF: Judge Agent Forest, a framework in which the judge …

Calibrate, Don't Curate: Label-Efficient Estimation from Noisy LLM Judges

2026-05-10 · Yanran Li arxiv

Multi-judge evaluation is increasingly used to assess LLMs and reward models, and the prevailing heuristic is to curate: keep the most accurate judges and discard weaker ones. We show that this heuristic can reverse when…

Revisiting Judge Decoding from First Principles via Training-Free Distributional Divergence

2026-01-08 · Shengyin Sun, Yiming Li, Renxi Liu, Weizhe Lin 외 arxiv

Judge Decoding accelerates LLM inference by relaxing the strict verification of Speculative Decoding, yet it typically relies on expensive and noisy supervision. In this work, we revisit this paradigm from first principl…