paper-with-me

홈 › Papers

Evaluating and Understanding Scheming Propensity in LLM Agents

2026-03-02 · Mia Hopman, Jannes Elstner, Maria Avramidou, Amritanshu Prasad, David Lindner arxiv

As frontier language models are increasingly deployed as autonomous agents pursuing complex, long-term objectives, there is increased risk of scheming: agents covertly pursuing misaligned goals. Prior work has focused on showing agents are capable of scheming, but their propensity to scheme in realistic scenarios remains underexplored. To understand when agents scheme, we decompose scheming incentives into agent factors and environmental factors. We develop realistic settings allowing us to systematically vary these factors, each with scheming opportunities for agents that pursue instrumentally convergent goals such as self-preservation, resource acquisition, and goal-guarding. We find only minimal instances of scheming despite high environmental incentives, and show this is unlikely due to evaluation awareness. While inserting adversarially-designed prompt snippets that encourage agency and goal-directedness into an agent's system prompt can induce high scheming rates, snippets used in real agent scaffolds rarely do. Surprisingly, in model organisms (Hubinger et al., 2023) built with these snippets, scheming behavior is remarkably brittle: removing a single tool can drop the scheming rate from 59% to 3%, and increasing oversight can raise rather than deter scheming by up to 25%. Our incentive decomposition enables systematic measurement of scheming propensity in settings relevant for deployment, which is necessary as agents are entrusted with increasingly consequential tasks.

📄 PDF Abstract BibTeX arXiv:2603.01608

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Scheming Ability in LLM-to-LLM Strategic Interactions

2025-10-11 · Thao Pham arxiv

As large language model (LLM) agents are deployed autonomously in diverse contexts, evaluating their capacity for strategic deception becomes crucial. While recent research has examined how AI systems scheme against huma…

SchemeArena: Factorized Stress Testing of Scheming in LLM Agents

2026-09-08 · Jie Ruan, Inderjeet Nair, Amy Liu, Muhammad Khalifa 외 hf

We study scheming in LLM agents, in which agents covertly pursue misaligned goals. Our focus is to understand how scheming arises from the interaction of key factors, such as instrumental goals, environmental affordances…

Stress Testing Deliberative Alignment for Anti-Scheming Training

2025-09-19 · Bronson Schoen, Evgenia Nitishinskaya, Mikita Balesni, Axel Højmark 외 arxiv

Highly capable AI systems could secretly pursue misaligned goals -- what we call "scheming". Because a scheming AI would deliberately try to hide its misaligned goals and actions, measuring and mitigating scheming requir…

Realistic honeypot evaluations for scheming propensity

2026-05-28 · Victoria Krakovna, David Lindner, Lewis Ho, Sebastian Farquhar 외 arxiv

We introduce scheming honeypot evaluations, a framework for testing whether models will pursue instrumental goals if given the opportunity. Our scheming honeypot evaluations take the form of coding tasks in Google's alig…

Frontier Models are Capable of In-context Scheming

2024-12-06 · Alexander Meinke, Bronson Schoen, Jérémy Scheurer, Mikita Balesni 외

Frontier models are increasingly trained and deployed as autonomous agent. One safety concern is that AI agents might covertly pursue misaligned goals, hiding their true capabilities and objectives - also known as schemi…