paper-with-me

홈 › Papers

Alignment Risks from Capability-Seeking RL Training

2026-02-12 · Yujun Zhou, Yue Huang, Han Bao, Kehan Guo, Zhenwen Liang, Pin-Yu Chen, Tian Gao, Werner Geyer, Nuno Moniz, Nitesh V Chawla, Xiangliang Zhang arxiv

While most AI alignment research focuses on preventing models from generating explicitly harmful content, a more subtle risk arises from capability-seeking RL training in vulnerable environments. We investigate whether language models, when trained with reinforcement learning (RL) in environments with implicit loopholes, can learn to exploit these flaws to maximize reward, even without being explicitly instructed to do so. To test this, we design a suite of four diverse "vulnerability games," each presenting a structural vulnerability related to context-conditional compliance, proxy metrics, reward tampering, and self-evaluation. Our experiments show that models often learn to exploit these vulnerabilities, discovering opportunistic strategies that increase reward while sometimes preserving or even improving standard task-performance metrics. More critically, we find that these exploitative strategies are not always narrow "tricks": they can transfer in structured but limited ways, propagate from a capable teacher model to other student models through SFT, and in several cases remain more persistent when learned through RL than when distilled through SFT. Our findings show that alignment risks from capability-seeking RL training can be difficult to detect with standard performance monitoring, suggesting that future AI safety work should extend beyond content moderation to auditing and securing training environments, reward mechanisms, and evaluation channels. Code is available at https://github.com/YujunZhou/Capability-seeking-RL-risk.

📄 PDF Abstract BibTeX arXiv:2602.12124

Code (0)

등록된 구현이 없습니다.

Tasks

Reinforcement Learning

Similar Papers 제목 키워드 기반

A Review of the Evidence for Existential Risk from AI via Misaligned Power-Seeking

2023-10-27 · Rose Hadshar

Rapid advancements in artificial intelligence (AI) have sparked growing concerns among experts, policymakers, and world leaders regarding the potential for increasingly advanced AI systems to pose existential risks. This…

OpenEval: Benchmarking Chinese LLMs across Capability, Alignment and Safety

2024-03-18 · Chuang Liu, Linhao Yu, Jiaxuan Li, Renren Jin 외

The rapid development of Chinese large language models (LLMs) poses big challenges for efficient LLM evaluation. While current initiatives have introduced new benchmarks or evaluation platforms for assessing Chinese LLMs…

BenchmarkingMathematical Reasoning

Understanding and Avoiding AI Failures: A Practical Guide

2021-04-22 · Heather M. Williams, Roman V. Yampolskiy

As AI technologies increase in capability and ubiquity, AI accidents are becoming more common. Based on normal accident theory, high reliability theory, and open systems theory, we create a framework for understanding th…

AgentMisalignment: Measuring the Propensity for Misaligned Behaviour in LLM-Based Agents

2025-06-04 · Akshat Naik, Patrick Quinn, Guillermo Bosch, Emma Gouné 외

As Large Language Model (LLM) agents become more widespread, associated misalignment risks increase. Prior work has examined agents' ability to enact misaligned behaviour (misalignment capability) and their compliance wi…

Large Language ModelPrompt Engineering

Model evaluation for extreme risks

2023-05-24 · Toby Shevlane, Sebastian Farquhar, Ben Garfinkel, Mary Phuong 외

Current approaches to building general-purpose AI systems tend to produce systems with both beneficial and harmful capabilities. Further progress in AI development could lead to capabilities that pose extreme risks, such…

model