paper-with-me

홈 › Papers

Breaking the Safety-Capability Tradeoff: Reinforcement Learning with Verifiable Rewards Maintains Safety Guardrails in LLMs

2025-11-26 · Dongkyu Derek Cho, Huan Song, Arijit Ghosh Chowdhury, Haotian An, Yawei Wang, Rohit Thekkanal, Negin Sokhandan, Sharlina Keshava, Hannah Marlowe arxiv

Fine-tuning large language models (LLMs) for downstream tasks typically exhibit a fundamental safety-capability tradeoff, where improving task performance degrades safety alignment even on benign datasets. This degradation persists across standard approaches including supervised finetuning (SFT) and reinforcement learning from human feedback (RLHF). While reinforcement learning with verifiable rewards (RLVR) has emerged as a promising alternative that optimizes models on objectively measurable tasks, its safety implications remain unexplored. We present the first comprehensive theoretical and empirical analysis of safety properties in RLVR. Theoretically, we derive upper bounds on safety drift under KL-constrained optimization and prove conditions under which safety degradation is eliminated. Empirically, we conduct extensive experiments across five adversarial safety benchmarks, demonstrating that RLVR can simultaneously enhance reasoning capabilities while maintaining or improving safety guardrails. Our comprehensive ablation studies examine the effects of optimization algorithms, model scale, and task domains. Our findings challenge the prevailing assumption of an inevitable safety capability trade-off, and establish that a specific training methodology can achieve both objectives simultaneously, providing insights for the safe deployment of reasoning-capable LLMs.

📄 PDF Abstract BibTeX arXiv:2511.21050

Code (0)

등록된 구현이 없습니다.

Tasks

Reinforcement Learning

Similar Papers 제목 키워드 기반

AlphaAlign: Incentivizing Safety Alignment with Extremely Simplified Reinforcement Learning

2025-07-20 · Yi Zhang, An Zhang, XiuYu Zhang, Leheng Sheng 외 arxiv

Large language models (LLMs), despite possessing latent safety understanding from their vast pretraining data, remain vulnerable to generating harmful content and exhibit issues such as over-refusal and utility degradati…

Reinforcement Learning

What Is the Alignment Tax?

2026-02-09 · Robin Young arxiv

The alignment tax is widely discussed but has not been formally characterized. We provide a geometric theory of the alignment tax in representation space. Under linear representation assumptions, we define the alignment …

David vs. Goliath: Verifiable Agent-to-Agent Jailbreaking via Reinforcement Learning

2026-02-02 · Samuel Nellessen, Tal Kachman arxiv

The evolution of large language models into autonomous agents introduces adversarial failures that exploit legitimate tool privileges, transforming safety evaluation in tool-augmented environments from a subjective NLP t…

Reinforcement Learning

SDGO: Self-Discrimination-Guided Optimization for Consistent Safety in Large Language Models

2025-08-21 · Peng Ding, Wen Sun, Dailin Li, Wei Zou 외 arxiv

Large Language Models (LLMs) excel at various natural language processing tasks but remain vulnerable to jailbreaking attacks that induce harmful content generation. In this paper, we reveal a critical safety inconsisten…

Reinforcement Learning

CARE-RL: Capability-Aware Reinforcement Learning for Mitigating Cross-Domain Conflicts

2026-05-30 · Rui Zhang, Xinle Wu, Yao Lu arxiv

Reinforcement learning (RL) with verifiable rewards has achieved strong progress in reasoning-oriented LLMs, but extending it to multi-domain RL remains challenging due to reward unreliability in non-verifiable tasks and…

Reinforcement Learning