paper-with-me

홈 › Papers

Sycophancy to Subterfuge: Investigating Reward-Tampering in Large Language Models

2024-06-14 · Carson Denison, Monte MacDiarmid, Fazl Barez, David Duvenaud, Shauna Kravec, Samuel Marks, Nicholas Schiefer, Ryan Soklaski, Alex Tamkin, Jared Kaplan, Buck Shlegeris, Samuel R. Bowman, Ethan Perez, Evan Hubinger

In reinforcement learning, specification gaming occurs when AI systems learn undesired behaviors that are highly rewarded due to misspecified training goals. Specification gaming can range from simple behaviors like sycophancy to sophisticated and pernicious behaviors like reward-tampering, where a model directly modifies its own reward mechanism. However, these more pernicious behaviors may be too complex to be discovered via exploration. In this paper, we study whether Large Language Model (LLM) assistants which find easily discovered forms of specification gaming will generalize to perform rarer and more blatant forms, up to and including reward-tampering. We construct a curriculum of increasingly sophisticated gameable environments and find that training on early-curriculum environments leads to more specification gaming on remaining environments. Strikingly, a small but non-negligible proportion of the time, LLM assistants trained on the full curriculum generalize zero-shot to directly rewriting their own reward function. Retraining an LLM not to game early-curriculum environments mitigates, but does not eliminate, reward-tampering in later environments. Moreover, adding harmlessness training to our gameable environments does not prevent reward-tampering. These results demonstrate that LLMs can generalize from common forms of specification gaming to more pernicious reward tampering and that such behavior may be nontrivial to remove.

📄 PDF Abstract BibTeX arXiv:2406.10162

Code (1)

anthropics/sycophancy-to-subterfuge-paper 공식 구현

Tasks

Language ModellingLarge Language Model

Similar Papers 제목 키워드 기반

Reward Tampering Problems and Solutions in Reinforcement Learning: A Causal Influence Diagram Perspective

2019-08-13 · Tom Everitt, Marcus Hutter, Ramana Kumar, Victoria Krakovna

Can humans get arbitrarily capable reinforcement learning (RL) agents to do their bidding? Or will sufficiently capable RL agents always find ways to bypass their intended objectives by shortcutting their reward signal? …

reinforcement-learningReinforcement LearningReinforcement Learning (RL)

Pressure, What Pressure? Sycophancy Disentanglement in Language Models via Reward Decomposition

2026-04-07 · Muhammad Ahmed Mohsin, Ahsan Bilal, Muhammad Umer, Emily Fox arxiv

Large language models exhibit sycophancy, the tendency to shift their stated positions toward perceived user preferences or authority cues regardless of evidence. Standard alignment methods fail to correct this because s…

Linear Probe Penalties Reduce LLM Sycophancy

2024-12-01 · Henry Papadatos, Rachel Freedman

Large language models (LLMs) are often sycophantic, prioritizing agreement with their users over accurate or objective statements. This problematic behavior becomes more pronounced during reinforcement learning from huma…

Investigating the Influence of Language on Sycophantic Behavior of Multilingual LLMs

2026-03-29 · Bayan Abdullah Aldahlawi, A. B. M. Ashikur Rahman, Irfan Ahmad arxiv

Large language models (LLMs) have achieved strong performance across a wide range of tasks, but they are also prone to sycophancy, the tendency to agree with user statements regardless of validity. Previous research has …

REALab: An Embedded Perspective on Tampering

2020-11-17 · Ramana Kumar, Jonathan Uesato, Richard Ngo, Tom Everitt 외

This paper describes REALab, a platform for embedded agency research in reinforcement learning (RL). REALab is designed to model the structure of tampering problems that may arise in real-world deployments of RL. Standar…

Reinforcement Learning (RL)