paper-with-me

홈 › Papers

Auditing Sabotage Bench: A Benchmark for Detecting and Fixing Research Sabotage in ML Codebases

2026-04-17 · Eric Gan, Aryan Bhatt, Buck Shlegeris, Julian Stastny, Vivek Hebbar arxiv

As AI systems are increasingly used to conduct research autonomously, misaligned systems could introduce subtle flaws that produce misleading results while evading detection. We introduce Auditing Sabotage Bench, a benchmark for evaluating the ability of auditors to detect and fix sabotage in ML research codebases. Our benchmark consists of 9 ML research codebases with sabotaged variants that produce qualitatively different experimental results. Each sabotage modifies implementation details, such as hyperparameters, training data, or evaluation code, while preserving the high-level methodology described in the paper. We evaluated frontier LLMs and LLM-assisted human auditors on our benchmark and found that both struggled to reliably detect and fix sabotage: the best performance was an AUROC of 0.77 and a top-1 fix rate of 42%, achieved by Gemini 3.1 Pro. We also tested LLMs as red teamers and found that LLM-generated sabotages were weaker than human-generated ones but still sometimes evaded same-capability LLM auditors. We release the benchmark to support research on monitoring and auditing techniques for AI-conducted research.

📄 PDF Abstract BibTeX arXiv:2604.16286

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Gram: Assessing sabotage propensities via automated alignment auditing

2026-05-28 · David Lindner, Victoria Krakovna, Sebastian Farquhar arxiv

We introduce Gram, an automated alignment auditing framework to assess the propensity of AI agents to engage in sabotage. We evaluate Gemini models across 17 simulated agentic deployment scenarios that incentivize sabota…

CTRL-ALT-DECEIT: Sabotage Evaluations for Automated AI R&D

2025-11-13 · Francis Rhys Ward, Teun van der Weij, Hanna Gábor, Sam Martin 외 arxiv

AI systems are increasingly able to autonomously conduct realistic software engineering tasks, and may soon be deployed to automate machine learning (ML) R&D itself. Frontier AI systems may be deployed in safety-critical…

When One Modality Sabotages the Others: A Diagnostic Lens on Multimodal Reasoning

2025-11-04 · Chenyu Zhang, Minsol Kim, Shohreh Ghorbani, Jingyao Wu 외 arxiv

Despite rapid growth in multimodal large language models (MLLMs), their reasoning traces remain opaque: it is often unclear which modality drives a prediction, how conflicts are resolved, or when one stream dominates. In…

Multimodal Emotion RecognitionMultimodal Reasoning

Evaluating whether AI models would sabotage AI safety research

2026-04-27 · Robert Kirk, Alexandra Souly, Kai Fronsdal, Abby D'Cruz 외 arxiv

We evaluate the propensity of frontier models to sabotage or refuse to assist with safety research when deployed as AI research agents within a frontier AI company. We apply two complementary evaluations to four Claude m…

Detecting Safety Violations Across Many Agent Traces

2026-04-13 · Adam Stein, Davis Brown, Hamed Hassani, Mayur Naik 외 arxiv

To identify safety violations, auditors often search over large sets of agent traces. This search is difficult because failures are often rare, complex, and sometimes even adversarially hidden and only detectable when mu…