paper-with-me

홈 › Papers

Judge, Retrieve, or Abstain: Uncertainty-Guarded LLM Judging with Provable Risk Guarantees

2026-08-18 · Sher Badshah, Ali Emami, Hassan Sajjad arxiv

Using LLMs as judges has become standard practice for evaluating model outputs at scale. This is particularly common for subjective, open-ended tasks such as assessing helpfulness or alignment, where no single reference answer exists. However, objective tasks introduce a distinct reliability challenge for reference-free LLM judging. In the absence of a reference answer, the judge evaluates factual correctness either through its parametric knowledge or through tool augmentation. Although the former enables efficient evaluation, the judge may hallucinate or lack sufficient evidence for its verdict. Conversely, tool augmentation can provide additional evidence but introduces extra computational cost and requires an appropriate mechanism to determine when and how that evidence should be used reliably. More importantly, neither approach alone provides formal control over the risk of accepted verdicts or guarantees their reliability at a specified level. We propose a risk-controlled framework that calibrates uncertainty thresholds on a held-out set so that the false discovery rate among accepted verdicts remains below a user-specified level~$α$ with high probability, using finite-sample Clopper--Pearson intervals. When the parametric mode is not sufficiently confident, the instance is routed to a retrieval-augmented mode, where the judge gathers web evidence and re-evaluates the instance under a second calibrated threshold. The finite-sample guarantee carries over to this two-threshold routing without additional assumptions. Across open-domain QA benchmarks and judges of varying scales, the framework maintains the target error rate while achieving substantially higher coverage than single-mode baselines.

📄 PDF Abstract BibTeX arXiv:2608.17994

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

SCOPE: Selective Conformal Optimized Pairwise LLM Judging

2026-02-13 · Sher Badshah, Ali Emami, Hassan Sajjad arxiv

Large language models (LLMs) are increasingly used as scalable judges in pairwise evaluation, but they remain prone to miscalibration and biases. We propose \textsc{Scope} (Selective Conformal Optimized Pairwise Evaluati…

S2G-RAG: Structured Sufficiency and Gap Judging for Iterative Retrieval-Augmented QA

2026-04-26 · Minghan Li, Junjie Zou, Xinxuan Lv, Chao Zhang 외 arxiv

Retrieval-Augmented Generation (RAG) grounds language models in external evidence, but multi-hop question answering remains difficult because iterative pipelines must control what to retrieve next and when the available …

Multi-hop Question Answering

JudgeStealer: Extracting LLM Judging Capabilities across Evaluation Protocols

2026-08-27 · Chen Chen, Yaolin Chen, Xuehan Sun, Juan Lin 외 arxiv

Large language model (LLM) judges are increasingly used across various evaluation scenarios, making their judgment capabilities valuable intellectual property. However, black-box access exposes these capabilities to mode…

Model extraction

Bittensor Agent Arenas as a Trajectory Primitive: Distilling a Shopping Agent from ShoppingBench Subnet Traces

2026-06-08 · Shardul Bansal, Seth Schilbe, Jarrod Barnes arxiv

Small-model agentic post-training is bottlenecked less by the algorithm than by the trajectory substrate it consumes. Leading recipes (RLVR, group-relative RL, rejection-sampled re-SFT) all need multi-turn traces carryin…

Permutation-Consensus Listwise Judging for Robust Factuality Evaluation

2026-03-20 · Tianyi Huang, Nathan Huang, Justin Tang, Wenqian Chen 외 arxiv

Large language models (LLMs) are now widely used as judges, yet their decisions can change under presentation choices that should be irrelevant. We study one such source of instability: candidate-order sensitivity in lis…