paper-with-me

홈 › Papers

From Reward-Hack Activations to Agentic Risk States: Context-Calibrated Mechanistic Monitoring in LLM Agents

2026-06-04 · Patrick Wilhelm, Odej Kao arxiv

Language-model agents act through repeated cycles of observation, reasoning, and action selection, making safety monitoring depend on both internal model state and environment context. We study reward-hacking monitors in ReAct-style agents acting in Gameable ALFWorld and WebShop. Agents are instrumented with activation-based reward-hack scores, token-level entropy, and decision-context features. We find that adapters fine-tuned on \textit{School-of-Reward-Hacks} dataset can transfer reward-hack tendencies into agentic action selection, especially when the environment exposes proxy-reward affordances. However, mitigating such behavior cannot rely on activation dynamics alone. High reward-hack activation identifies a latent policy state, but does not necessarily imply an immediate exploit action. Across next-step prediction tasks, entropy and context-calibrated internal features improve risk estimation over reward-hack activation alone. Activation-direction steering further reduces proxy-exploit behavior in selected mixed-adapter regimes. Overall, our results support context-calibrated internal monitoring for agents: reward-hack activation identifies a latent policy state, while entropy and decision context help determine when that state becomes risky action.

📄 PDF Abstract BibTeX arXiv:2606.06223

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

RewardFlow: Topology-Aware Reward Propagation on State Graphs for Agentic RL with Large Language Models

2026-03-19 · Xiao Feng, Bo Han, Zhanke Zhou, Jiaqi Fan 외 arxiv

Reinforcement learning (RL) shows promise for enhancing LLM agentic reasoning, yet sparse terminal rewards hinder fine-grained optimization. Process reward modeling offers an alternative but incurs high computational cos…

Reinforcement LearningVisual Reasoning

The Obfuscation Atlas: Mapping Where Honesty Emerges in RLVR with Deception Probes

2026-02-17 · Mohammad Taufeeque, Stefan Heimersheim, Adam Gleave, Chris Cundy arxiv

Training against white-box deception detectors has been proposed as a way to make AI systems honest. However, such training risks models learning to obfuscate their deception to evade the detector. Prior work has studied…

Monitoring Emergent Reward Hacking During Generation via Internal Activations

2026-03-04 · Patrick Wilhelm, Thorsten Wittkopp, Odej Kao arxiv

Fine-tuned large language models can exhibit reward-hacking behavior arising from emergent misalignment, which is difficult to detect from final outputs alone. While prior work has studied reward hacking at the level of …

Monitoring Reasoning Models for Misbehavior and the Risks of Promoting Obfuscation

2025-03-14 · Bowen Baker, Joost Huizinga, Leo Gao, Zehao Dou 외

Mitigating reward hacking--where AI systems misbehave due to flaws or misspecifications in their learning objectives--remains a key challenge in constructing capable and aligned models. We show that we can monitor a fron…

Natural Emergent Misalignment from Reward Hacking in Production RL

2025-11-23 · Monte MacDiarmid, Benjamin Wright, Jonathan Uesato, Joe Benton 외 arxiv

We show that when large language models learn to reward hack on production RL environments, this can result in egregious emergent misalignment. We start with a pretrained model, impart knowledge of reward hacking strateg…