paper-with-me

홈 › Papers

Greed Is Learned: Visible Incentives as Reward-Hacking Triggers

2026-06-15 · Tong Che, Rui Wu arxiv

Deployed agents increasingly act with their reward proxy in view, such as a balance, score, or KPI dashboard. We show that reinforcement learning can make a policy \emph{addicted} to such a visible self-benefit channel. It chases the displayed payoff across held-out domains, sacrifices the true task to do so, and follows the channel wherever we rewrite it, while policies that never saw the channel stay honest. We call this \emph{reward-channel addiction} and study it in \emph{MoneyWorld}, a synthetic sandbox. The addiction can \emph{flip a model's safety alignment}: trained only on innocuous money tasks with no safety content, the model abandons the safe action it otherwise always takes whenever a dashboard pays for an unsafe one, and reverts to safe once the channel is hidden. This learned bribe replicates across model scales and families. Blindly optimizing super-capable, next-generation AI on KPIs or P\&L can be dangerous for alignment. \emph{Greed is learned} when following such a channel pays.

📄 PDF Abstract BibTeX arXiv:2606.16914

Code (0)

등록된 구현이 없습니다.

Tasks

Reinforcement Learning

Similar Papers 제목 키워드 기반

Proxy Reward Internalization and Mechanistic Exploitation: A Learned Precursor to Reward Hacking and Its Generalization

2026-06-08 · Mohammad Beigi, Ming Jin, Lifu Huang arxiv

Reward hacking is usually studied after it becomes visible, once a model earns high proxy reward while failing the intended task. We instead study what proxy RL teaches before that failure appears. We introduce Proxy Rew…

SpecBench: Measuring Reward Hacking in Long-Horizon Coding Agents

2026-05-20 · Bingchen Zhao, Dhruv Srikanth, Yuxiang Wu, Zhengyao Jiang arxiv

As long-horizon coding agents produce more code than any developer can review, oversight collapses onto a single surface: the automated test suite. Reward hacking naturally arises in this setup, as the agent optimizes fo…

Debate Training Reduces Reward Hacking in RLAIF

2026-08-18 · Zachary Kenton, Lili Janzer, Rory Greig, Tian Huey Teh 외 arxiv

We demonstrate that RL finetuning an LLM using debate, a two-player adversarial game between a generator and a critic adjudicated by a weaker LLM judge, reduces reward hacking compared to a reinforcement learning from AI…

Reinforcement Learning

Discovering Implicit Large Language Model Alignment Objectives

2026-02-17 · Edward Chen, Sanmi Koyejo, Carlos Guestrin arxiv

Large language model (LLM) alignment relies on complex reward signals that often obscure the specific behaviors being incentivized, creating critical risks of misalignment and reward hacking. Existing interpretation meth…

Extending MONA in Camera Dropbox: Reproduction, Learned Approval, and Design Implications for Reward-Hacking Mitigation

2026-03-31 · Nathan Heath arxiv

Myopic Optimization with Non-myopic Approval (MONA) mitigates multi-step reward hacking by restricting the agent's planning horizon while supplying far-sighted approval as a training signal~\cite{farquhar2025mona}. The o…