paper-with-me

Papers

When Judgment Becomes Noise: How Design Failures in LLM Judge Benchmarks Silently Undermine Validity

2025-09-24 · Benjamin Feuer, Chiung-Yi Tseng, Astitwa Sarthak Lathe, Oussama Elachqar, John P Dickerson arxiv

LLM-judged benchmarks are increasingly used to evaluate complex model behaviors, yet their design introduces failure modes absent in conventional ground-truth based benchmarks. We argue that without tight objectives and verifiable constructions, benchmark rankings can produce high-confidence rankings that are in fact largely noise. We introduce two mechanisms to diagnose these issues. Schematic adherence quantifies how much of a judge's overall verdict is explained by the explicit evaluation schema, revealing unexplained variance when judges deviate from their own rubric. Psychometric validity aggregates internal consistency and discriminant validity signals to quantify irreducible uncertainty in any benchmarking run. Applying these tools to Arena-Hard Auto, we find severe schema incoherence and factor collapse across popular judges: for example, unexplained variance exceeding 90 percent for DeepSeek-R1-32B and factor correlations above 0.93 for most criteria. We also show that the ELO-style aggregation used by Arena-Hard Auto collapses and masks genuine ranking uncertainty. Our results highlight design failures that undermine validity and offer actionable principles for building better-scoped, reliability-aware LLM-judged benchmarks. We released our code and dataset at https://github.com/penfever/judgment-to-noise

📄 PDF Abstract BibTeX arXiv:2509.20293

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Transitive Expert Error and Routing Problems in Complex AI Systems

2026-01-07 · Forest Mars arxiv

Domain expertise enhances judgment within boundaries but creates systematic vulnerabilities specifically at borders. We term this Transitive Expert Error (TEE), distinct from Dunning-Kruger effects, requiring calibrated …

AI Alignment From Social Choice Perspectives

2026-06-19 · Daniel Halpern, Evi Micha, Ariel D. Procaccia, Benjamin Schiffer 외 arxiv

Alignment from human feedback uses human judgments about model outputs to steer the behavior of language models after pretraining. When those judgments reflect conflicting views of desirable behavior, the learned objecti…

AI Failures: A Review of Underlying Issues

2020-07-18 · Debarag Narayan Banerjee, Sasanka Sekhar Chanda

Instances of Artificial Intelligence (AI) systems failing to deliver consistent, satisfactory performance are legion. We investigate why AI failures occur. We address only a narrow subset of the broader field of AI Safet…

Cheap Code, Costly Judgment: A Case Study on Governable Agentic Software Engineering

2026-07-01 · James C. Davis, Paschal C. Amusuo, Tanmay Singla, Berk Çakar 외 arxiv

Generative AI is shifting software engineering from a practice organized around scarce implementation effort toward one organized around abundant, low-cost code production. This shift changes the central engineering prob…

Cascading Waves of Fluctuation in Time-delay Multi-agent Rendezvous

2023-03-15 · Guangyi Liu, Vivek Pandey, Christoforos Somarakis, Nader Motee

We develop a framework to assess the risk of cascading failures when a team of agents aims to rendezvous in time in the presence of exogenous noise and communication time-delay. The notion of value-at-risk (VaR) measure …