paper-with-me

Papers

Safety Training Modulates Harmful Misalignment Under On-Policy RL, But Direction Depends on Environment Design

2026-04-14 · Leon Eshuijs, Shihan Wang, Antske Fokkens arxiv

Specification gaming under Reinforcement Learning (RL) is known to cause LLMs to develop sycophantic, manipulative, or deceptive behavior, yet the conditions under which this occurs remain unclear. We train 11 instruction-tuned LLMs (0.5B--14B) with on-policy RL across 3 environments and find that model size acts as a safety buffer in some environments but enables greater harmful exploitation in others. Controlled ablations trace this reversal to environment-specific features such as role framing and implicit gameability cues. We further show that most safety benchmarks do not predict RL-induced misalignment, except in the case of Sycophancy scores when the exploit relies on inferring the user's preference. Finally, we find that on-policy RL preserves a safety buffer inherent in the model's own generation distribution, one that is bypassed during off-policy settings.

📄 PDF Abstract BibTeX arXiv:2604.12500

Code (0)

등록된 구현이 없습니다.

Tasks

Reinforcement Learning

Similar Papers 제목 키워드 기반

Narrow Fine-Tuning Erodes Safety Alignment in Vision-Language Agents

2026-02-18 · Idhant Gulati, Shivam Raval arxiv

Lifelong multimodal agents must continuously adapt to new tasks through post-training, but this creates a fundamental tension between acquiring capabilities and preserving safety alignment. We demonstrate that fine-tunin…

Continual Learning

Understanding Emergent Misalignment via Feature Superposition Geometry

2026-04-07 · Gouki Minegishi, Hiroki Furuta, Takeshi Kojima, Yusuke Iwasawa 외 arxiv

Emergent misalignment, where fine-tuning on narrow, non-harmful tasks induces harmful behaviors, poses a key challenge for AI safety in LLMs. Despite growing empirical evidence, its underlying mechanism remains unclear. …

When Thinking Backfires: Mechanistic Insights Into Reasoning-Induced Misalignment

2025-08-30 · Hanqi Yan, Hainiu Xu, Siya Qi, Shu Yang 외 arxiv

With the growing accessibility and wide adoption of large language models, concerns about their safety and alignment with human values have become paramount. In this paper, we identify a concerning phenomenon: Reasoning-…

The Emotional Baby Is Truly Deadly: Does your Multimodal Large Reasoning Model Have Emotional Flattery towards Humans?

2025-08-06 · Yuan Xun, Xiaojun Jia, Xinwei Liu, Hua Zhang arxiv

We observe that MLRMs oriented toward human-centric service are highly susceptible to user emotional cues during the deep-thinking stage, often overriding safety protocols or built-in safety checks under high emotional i…

Semantic Containment as a Fundamental Property of Emergent Misalignment

2026-02-02 · Rohan Saxena arxiv

Fine-tuning language models on narrowly harmful data causes emergent misalignment (EM) -- behavioral failures extending far beyond training distributions. Recent work demonstrates compartmentalization of misalignment beh…