paper-with-me

홈 › Papers

ResearchArena: Evaluating Sabotage and Monitoring in Automated AI R&D

2026-07-21 · Lena Libon, Ben Rank, Jehyeok Yeon, David Schmotz, Jeremy Qin, Daniel Donnelly, Derck Prinzhorn, Maksym Andriushchenko arxiv

As AI agents begin to automate AI R&D, we need ways to assess whether their outputs are safe to deploy, even when the agents themselves may be untrusted. AI control offers one such approach: rather than trusting the agent, it treats it as a potential adversary and uses a monitor to detect covert sabotage before deployment. We evaluate AI control for automated AI R&D with ResearchArena, a framework spanning four long-horizon tasks: safety post-training, capabilities post-training, CUDA-kernel optimization, and inference-server optimization. Because the deliverable in AI R&D is an artifact that will be deployed and run, we pair each main task with two kinds of hidden side task: an embedded side task that sabotages the submitted model, adapter, kernel, or server itself, and an independent side task that takes a covert out-of-scope action in the sandbox while leaving the artifact honest. We evaluate frontier agents at both sabotage and monitoring, across four monitor types that vary along two axes: whether the monitor may execute and probe the produced artifact, and whether it sees the agent's chain-of-thought. We find that sabotage hidden in the training data is the hardest to catch, flagged fewer than half the time. Letting the monitor run experiments on the artifact, rather than only read the trajectory, helps, but it is not enough: monitors still miss embedded sabotage by inspecting only the surface, by noticing the anomaly but explaining it away, or by probing the artifact with the wrong test. We release ResearchArena as a modular framework for evaluating sabotage and control in automated AI R&D.

📄 PDF Abstract BibTeX arXiv:2607.19321

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

SHADE-Arena: Evaluating Sabotage and Monitoring in LLM Agents

2025-06-17 · Jonathan Kutasov, Yuqi Sun, Paul Colognese, Teun van der Weij 외

As Large Language Models (LLMs) are increasingly deployed as autonomous agents in complex and long horizon settings, it is critical to evaluate their ability to sabotage users by pursuing hidden objectives. We study the …

Auditing Sabotage Bench: A Benchmark for Detecting and Fixing Research Sabotage in ML Codebases

2026-04-17 · Eric Gan, Aryan Bhatt, Buck Shlegeris, Julian Stastny 외 arxiv

As AI systems are increasingly used to conduct research autonomously, misaligned systems could introduce subtle flaws that produce misleading results while evading detection. We introduce Auditing Sabotage Bench, a bench…

CTRL-ALT-DECEIT: Sabotage Evaluations for Automated AI R&D

2025-11-13 · Francis Rhys Ward, Teun van der Weij, Hanna Gábor, Sam Martin 외 arxiv

AI systems are increasingly able to autonomously conduct realistic software engineering tasks, and may soon be deployed to automate machine learning (ML) R&D itself. Frontier AI systems may be deployed in safety-critical…

How does information access affect LLM monitors' ability to detect sabotage?

2026-01-28 · Rauno Arike, Raja Mehta Moreno, Rohan Subramani, Shubhorup Biswas 외 arxiv

Frontier language model agents can exhibit misaligned behaviors, including deception, exploiting reward hacks, and pursuing hidden objectives. To control potentially misaligned agents, we can use LLMs themselves to monit…

Gram: Assessing sabotage propensities via automated alignment auditing

2026-05-28 · David Lindner, Victoria Krakovna, Sebastian Farquhar arxiv

We introduce Gram, an automated alignment auditing framework to assess the propensity of AI agents to engage in sabotage. We evaluate Gemini models across 17 simulated agentic deployment scenarios that incentivize sabota…