paper-with-me

홈 › Papers

Gram: Assessing sabotage propensities via automated alignment auditing

2026-05-28 · David Lindner, Victoria Krakovna, Sebastian Farquhar arxiv

We introduce Gram, an automated alignment auditing framework to assess the propensity of AI agents to engage in sabotage. We evaluate Gemini models across 17 simulated agentic deployment scenarios that incentivize sabotage. We find Gemini models misbehave in about 2-3% of our simulated trajectories. Many of these cases are explained by "overeagerness" in Gemini models resulting in both excessive role-playing and goal-seeking behavior. In contrast to other alignment auditing approaches, Gram is designed to specifically evaluate misalignment and intentional sabotage in agentic coding and research agents. We additionally introduce an experimental investigator agent pipeline which enables fine-grained targeted experiments to identify the drivers of misbehavior. We find that increasing realism of environments and removing nudges to misbehave tends to reduce sabotage rates close to zero.

📄 PDF Abstract BibTeX arXiv:2605.30322

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

ResearchArena: Evaluating Sabotage and Monitoring in Automated AI R&D

2026-07-21 · Lena Libon, Ben Rank, Jehyeok Yeon, David Schmotz 외 arxiv

As AI agents begin to automate AI R&D, we need ways to assess whether their outputs are safe to deploy, even when the agents themselves may be untrusted. AI control offers one such approach: rather than trusting the agen…

UK AISI Alignment Evaluation Case-Study

2026-04-01 · Alexandra Souly, Robert Kirk, Jacob Merizian, Abby D'Cruz 외 arxiv

This technical report presents methods developed by the UK AI Security Institute for assessing whether advanced AI systems reliably follow intended goals. Specifically, we evaluate whether frontier models sabotage safety…

CTRL-ALT-DECEIT: Sabotage Evaluations for Automated AI R&D

2025-11-13 · Francis Rhys Ward, Teun van der Weij, Hanna Gábor, Sam Martin 외 arxiv

AI systems are increasingly able to autonomously conduct realistic software engineering tasks, and may soon be deployed to automate machine learning (ML) R&D itself. Frontier AI systems may be deployed in safety-critical…

Automated alignment is harder than you think

2026-05-07 · Aleksandr Bowkis, Marie Davidsen Buhl, Jacob Pfau, Geoffrey Irving arxiv

A leading proposal for aligning artificial superintelligence (ASI) is to use AI agents to automate an increasing fraction of alignment research as capabilities improve. We argue that, even when research agents are not sc…

Evaluating whether AI models would sabotage AI safety research

2026-04-27 · Robert Kirk, Alexandra Souly, Kai Fronsdal, Abby D'Cruz 외 arxiv

We evaluate the propensity of frontier models to sabotage or refuse to assist with safety research when deployed as AI research agents within a frontier AI company. We apply two complementary evaluations to four Claude m…