paper-with-me

Papers

Auditing the Audit: Five Failure Modes in Benchmark-Validity Audits

2026-07-01 · Yanhang Li, Zhichao Fan, Zexin Zhuang arxiv

Governance frameworks ask AI providers and auditors for documented evaluation evidence, and perturbation-based construct-validity audits are a common form of that evidence. We argue the audits are themselves fragile: their conclusions can be silently manufactured by implementation details that readers cannot see in the reported numbers. We name five classes of pipeline failure and demonstrate each in a self-audit over safety benchmarks and open-weight instruction-tuned models. Under a unified six-point due-diligence gate, every cell lands in a non-confirmatory bucket, and no cell reaches confirmatory. The evidence here is a single two-model, five-benchmark case study, and F1--F5 is an illustrative, deliberately non-exhaustive starting taxonomy -- not a comprehensive partition of audit failures. We position the gate as a withholding and disclosure protocol for assurance-grade evidence, supplementary to (not a replacement for) classical construct-validity evidence, and not as a route to benchmark-validity verdicts.

📄 PDF Abstract BibTeX arXiv:2607.02586

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Beyond Final Answers: Auditing Trajectory-Level Hallucinations in Multi-Agent Industrial Workflows

2026-05-22 · Harshada Badave, Santosh Borse, Andrea Gomez, Harshitha Narahari 외 arxiv

Large Language Models (LLMs) are increasingly deployed as autonomous agents that reason, use tools, and act over multiple steps. Yet most hallucination benchmarks still evaluate only the final output, missing failures th…

Differences That Matter: Auditing Models for Capability Gap Discovery and Rectification

2025-12-18 · Qihao Liu, Chengzhi Mao, Yaojie Liu, Alan Yuille 외 arxiv

Conventional evaluation methods for multimodal LLMs (MLLMs) lack interpretability and are often insufficient to fully disclose significant capability gaps across models. To address this, we introduce AuditDM, an automate…

Reinforcement Learning

An autonomous agent for auditing and improving the reliability of clinical AI models

2025-07-08 · Lukas Kuhn, Florian Buettner arxiv

The deployment of AI models in clinical practice faces a critical challenge: models achieving expert-level performance on benchmarks can fail catastrophically when confronted with real-world variations in medical imaging…

The Reliability Gap in Benchmark Auditing: Distribution Shift and Scale as Failure Modes of Contamination Detection

2026-06-02 · Wojciech Zarzecki, Jan Dubiński, Sebastian Cygert arxiv

Benchmark contamination, where evaluation examples appear in a model's training data, threatens the validity of LLM assessment. Statistical tools for detecting training-data membership exist, but have been validated almo…

Supporting Human-AI Collaboration in Auditing LLMs with LLMs

2023-04-19 · Charvi Rastogi, Marco Tulio Ribeiro, Nicholas King, Harsha Nori 외

Large language models are becoming increasingly pervasive and ubiquitous in society via deployment in sociotechnical systems. Yet these language models, be it for classification or generation, have been shown to be biase…

Language ModellingLarge Language ModelSentiment Analysis