paper-with-me

홈 › Papers

Mitigating the Safety-utility Trade-off in LLM Alignment via Adaptive Safe Context Learning

2026-02-14 · Yanbo Wang, Minzheng Wang, Jian Liang, Lu Wang, Yongcan Yu, Ran He arxiv

While reasoning models have achieved remarkable success in complex reasoning tasks, their increasing power necessitates stringent safety measures. For safety alignment, the core challenge lies in the inherent trade-off between safety and utility. However, prevailing alignment strategies typically construct CoT training data with explicit safety rules via context distillation. This approach inadvertently limits reasoning capabilities by creating a rigid association between rule memorization and refusal. To mitigate the safety-utility trade-off, we propose the Adaptive Safe Context Learning~(ASCL) framework to improve the reasoning given proper context. ASCL formulates safety alignment as a multi-turn tool-use process, empowering the model to autonomously decide when to consult safety rules and how to generate the ongoing reasoning. Furthermore, to counteract the preference for rule consultation during RL, we introduce Inverse Frequency Policy Optimization~(IFPO) to rebalance advantage estimates. By decoupling rule retrieval and subsequent reasoning, our method achieves higher overall performance compared to baselines. Our code is publicly available at https://github.com/ybwang119/ASCL.

📄 PDF Abstract BibTeX arXiv:2602.13562

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Safety Alignment as Continual Learning: Mitigating the Alignment Tax via Orthogonal Gradient Projection

2026-02-08 · Guanglong Sun, Siyuan Zhang, Liyuan Wang, Jun Zhu 외 arxiv

Safety post-training can improve the harmfulness and policy compliance of Large Language Models (LLMs), but it may also reduce general utility, a phenomenon often described as the \emph{alignment tax}. We study this trad…

Continual Learning

The Unintended Trade-off of AI Alignment:Balancing Hallucination Mitigation and Safety in LLMs

2025-10-09 · Omar Mahmoud, Ali Khalil, Buddhika Laknath Semage, Thommen George Karimpanal 외 arxiv

Hallucination in large language models (LLMs) has been widely studied in recent years, with progress in both detection and mitigation aimed at improving truthfulness. Yet, a critical side effect remains largely overlooke…

EVA: Editing for Versatile Alignment against Jailbreaks

2026-05-14 · Yi Wang, Hongye Qiu, Yue Xu, Sibei Yang 외 arxiv

Large Language Models (LLMs) and Vision Language Models (VLMs) have demonstrated impressive capabilities but remain vulnerable to jailbreaking attacks, where adversaries exploit textual or visual triggers to bypass safet…

Beyond Static Alignment: Hierarchical Policy Control for LLM Safety via Risk-Aware Chain-of-Thought

2026-02-06 · Jianfeng Si, Lin Sun, Weihong Lin, Xiangzheng Zhang arxiv

Large Language Models (LLMs) face a fundamental safety-helpfulness trade-off due to static, one-size-fits-all safety policies that lack runtime controllabilityxf, making it difficult to tailor responses to diverse applic…

The Rise of Darkness: Safety-Utility Trade-Offs in Role-Playing Dialogue Agents

2025-02-28 · Yihong Tang, Kehai Chen, Xuefeng Bai, ZhengYu Niu 외

Large Language Models (LLMs) have made remarkable advances in role-playing dialogue agents, demonstrating their utility in character simulations. However, it remains challenging for these agents to balance character port…