paper-with-me

홈 › Papers

AgentMisalignment: Measuring the Propensity for Misaligned Behaviour in LLM-Based Agents

2025-06-04 · Akshat Naik, Patrick Quinn, Guillermo Bosch, Emma Gouné, Francisco Javier Campos Zabala, Jason Ross Brown, Edward James Young

As Large Language Model (LLM) agents become more widespread, associated misalignment risks increase. Prior work has examined agents' ability to enact misaligned behaviour (misalignment capability) and their compliance with harmful instructions (misuse propensity). However, the likelihood of agents attempting misaligned behaviours in real-world settings (misalignment propensity) remains poorly understood. We introduce a misalignment propensity benchmark, AgentMisalignment, consisting of a suite of realistic scenarios in which LLM agents have the opportunity to display misaligned behaviour. We organise our evaluations into subcategories of misaligned behaviours, including goal-guarding, resisting shutdown, sandbagging, and power-seeking. We report the performance of frontier models on our benchmark, observing higher misalignment on average when evaluating more capable models. Finally, we systematically vary agent personalities through different system prompts. We find that persona characteristics can dramatically and unpredictably influence misalignment tendencies -- occasionally far more than the choice of model itself -- highlighting the importance of careful system prompt engineering for deployed AI agents. Our work highlights the failure of current alignment methods to generalise to LLM agents, and underscores the need for further propensity evaluations as autonomous systems become more prevalent.

📄 PDF Abstract BibTeX arXiv:2506.04018

Code (0)

등록된 구현이 없습니다.

Tasks

Large Language ModelPrompt Engineering

Similar Papers 제목 키워드 기반

Propensity Inference: Environmental Contributors to LLM Behaviour

2026-04-22 · Olli Järviniemi, Oliver Makins, Jacob Merizian, Robert Kirk 외 arxiv

Motivated by loss of control risks from misaligned AI systems, we develop and apply methods for measuring language models' propensity for unsanctioned behaviour. We contribute three methodological improvements: analysing…

Instrumental Choices: Measuring the Propensity of LLM Agents to Pursue Instrumental Behaviors

2026-05-07 · Jonas Wiedermann-Möller, Leonard Dung, Maksym Andriushchenko arxiv

AI systems have become increasingly capable of dangerous behaviours in many domains. This raises the question: Do models sometimes choose to violate human instructions in order to perform behaviour that is more useful fo…

Capabilities Ain't All You Need: Measuring Propensities in AI

2026-02-20 · Daniel Romero-Alvarado, Fernando Martínez-Plumed, Lorenzo Pacchiardi, Hugo Save 외 arxiv

AI evaluation has primarily focused on measuring capabilities, with formal approaches inspired from Item Response Theory (IRT) being increasingly applied. Yet propensities - the tendencies of models to exhibit particular…

Evaluating and Understanding Scheming Propensity in LLM Agents

2026-03-02 · Mia Hopman, Jannes Elstner, Maria Avramidou, Amritanshu Prasad 외 arxiv

As frontier language models are increasingly deployed as autonomous agents pursuing complex, long-term objectives, there is increased risk of scheming: agents covertly pursuing misaligned goals. Prior work has focused on…

Stress Testing Deliberative Alignment for Anti-Scheming Training

2025-09-19 · Bronson Schoen, Evgenia Nitishinskaya, Mikita Balesni, Axel Højmark 외 arxiv

Highly capable AI systems could secretly pursue misaligned goals -- what we call "scheming". Because a scheming AI would deliberately try to hide its misaligned goals and actions, measuring and mitigating scheming requir…