paper-with-me

홈 › Papers

Hidden Measurement Error in LLM Pipelines Distorts Annotation, Evaluation, and Benchmarking

2026-04-13 · Solomon Messing arxiv

LLM evaluations drive which models get deployed, what safety standards get adopted, which research conclusions get published, and how projections of AI's labor-market impact get made. Yet standard confidence intervals ignore variability from judge model choice, model temperature, and prompt phrasing, producing under-coverage that worsens with more data. The omitted variance can shift results enough to reverse conclusions \citep{baumann2025llmhacking, huang2026dropping}; pipelines that fail to average over it leave the surface that ``benchmark hacking'' exploits \citep{singh2025leaderboard}. This paper decomposes LLM pipeline uncertainty into its sources, distinguishes variance that shrinks with more data from sensitivity to researcher design choices, and uses design-study projections to reduce total evaluation error (TEE). Across the demonstrations, naive standard errors are 40 - 60\% smaller than the TEE-corrected SE. Using Chatbot Arena data, we show naive 95\% CI coverage drops as $n$ grows while TEE-corrected coverage holds at 95\%, and TEE-guided pipelines restrict the benchmark gaming surface from 56 to 32 Elo ($K=27$), below the human-leaderboard baseline. We show further that a small pilot recovers honest CIs and projects which design changes most improve precision. Acting on those projections halves MMLU estimation error against the answer key at equivalent cost, and raises per-match agreement with human votes by 7.9 percentage points on Chatbot Arena.

📄 PDF Abstract BibTeX arXiv:2604.11581

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Position: AI for Science Should Treat Measurement-to-Dataset Pipelines as Inference Components

2026-05-23 · Ling Zhan, Xiaoyao Yu, Tao Jia arxiv

AI for Science (AI4Science) workflows often treat the released dataset as a fixed interface to the underlying system. However, in domains relying on \emph{indirect observation}, the learner observes a derivative represen…

Train What You Deploy:Token-Faithful Post-Training of a Production Coding

2026-09-04 · Cheng Li, Jiexiong Liu, Yixuan Chen, Chi Hong arxiv

Existing post-training pipelines for coding and terminal agents suffer severe token and control fidelity errors: simplified training environments mismatch production deployments, and offline token reconstruction from age…

Automated Benchmark Auditing for AI Agents and Large Language Models

2026-05-25 · Junlin Wang, Federico Bianchi, Shang Zhu, Fan Nie 외 arxiv

Modern AI benchmarks operate at a complexity that outpaces traditional verification methods. Tasks authored by domain experts often contain implicit assumptions, incomplete environment specifications, and brittle evaluat…

A2Eval: Agentic and Automated Evaluation for Embodied Brain

2026-02-02 · Shuai Zhang, Jiayu Hu, Zijie Chen, Zeyuan Ding 외 arxiv

Current embodied VLM evaluation relies on static, expert-defined, manually annotated benchmarks that exhibit severe redundancy and coverage imbalance. This labor intensive paradigm drains computational and annotation res…

Towards Fairness under Label Bias in Image Segmentation: Impact, Measurement and Mitigation

2026-05-07 · Aditya Parikh, Stella Frank, Sneha Das, Aasa Feragen arxiv

Labeled datasets reflect the biases of their annotation pipelines, which sometimes introduce label bias: group-conditional label errors that cause systematic performance disparities across demographic subgroups. Label bi…

Image Segmentation