paper-with-me

홈 › Papers

RAJ-PGA: Reasoning-Activated Jailbreak and Principle-Guided Alignment Framework for Large Reasoning Models

2025-08-18 · Jianhao Chen, Mayi Xu, Haoyang Chen, Xiaohu Li, Xiangyu Zhang, Jianjie Huang, Zheng Wang, Xiaochun Cao, Tieyun Qian arxiv

Large Reasoning Models (LRMs) face a distinct safety vulnerability: their internal reasoning chains may generate harmful content even when the final output appears benign. To address this overlooked risk, we first propose a novel attack paradigm, Reasoning-Activated Jailbreak (RAJ) via Concretization, which demonstrates that refining malicious prompts to be more specific can trigger step-by-step logical reasoning that overrides the model's safety protocols. To systematically mitigate this vulnerability, we further develop a scalable framework for constructing high-quality safety alignment datasets. This framework first leverages the RAJ attack to elicit challenging harmful reasoning chains from LRMs, then transforms these high-risk traces into safe, constructive, and educational responses through a tailored Principle-Guided Alignment (PGA) mechanism. Then, we introduce the PGA dataset, a verified alignment dataset containing 3,989 samples using our proposed method. Extensive experiments show that fine-tuning LRMs with PGA dataset significantly enhances model safety, achieving up to a 29.5% improvement in defense success rates across multiple jailbreak benchmarks. Critically, our approach not only defends against sophisticated reasoning-based attacks but also preserves, even enhances, the model's general reasoning capabilities. This work provides a scalable and effective pathway for safety alignment in reasoning-intensive AI systems, addressing the core trade-off between safety and functional performance.

📄 PDF Abstract BibTeX arXiv:2508.12897

Code (0)

등록된 구현이 없습니다.

Tasks

Logical Reasoning

Similar Papers 제목 키워드 기반

ProMoral-Bench: Evaluating Prompting Strategies for Moral Reasoning and Safety in LLMs

2026-02-05 · Rohan Subramanian Thomas, Shikhar Shiromani, Abdullah Chaudhry, Ruizhe Li 외 arxiv

Prompt design significantly impacts the moral competence and safety alignment of large language models (LLMs), yet empirical comparisons remain fragmented across datasets and models.We introduce ProMoral-Bench, a unified…

Prompt Engineering

RSafe: Incentivizing proactive reasoning to build robust and adaptive LLM safeguards

2025-06-09 · Jingnan Zheng, Xiangtian Ji, Yijun Lu, Chenhang Cui 외

Large Language Models (LLMs) continue to exhibit vulnerabilities despite deliberate safety alignment efforts, posing significant risks to users and society. To safeguard against the risk of policy-violating content, syst…

Safety Alignment

Alignment-Weighted DPO: A principled reasoning approach to improve safety alignment

2026-02-24 · Mengxuan Hu, Vivek V. Datla, Anoop Kumar, Zihan Guan 외 arxiv

Recent advances in alignment techniques such as Supervised Fine-Tuning (SFT), Reinforcement Learning from Human Feedback (RLHF), and Direct Preference Optimization (DPO) have improved the safety of large language models …

Reinforcement Learning

ReFrame: Evidence-Guided Test-Time Safety Alignment in Multimodal Large Language Models

2026-08-21 · Wenzheng Jiang, Xuankun Rong, Yuanzhao Zhai, Dawei Feng 외 arxiv

While multimodal large language models (MLLMs) extend model capabilities beyond text, they also make safety alignment increasingly challenging. Multimodal safety alignment methods must address cross-modal jailbreaks, saf…

STAR-S: Improving Safety Alignment through Self-Taught Reasoning on Safety Rules

2026-01-07 · Di Wu, Yanyan Zhao, Xin Lu, Mingzhe Li 외 arxiv

Defending against jailbreak attacks is crucial for the safe deployment of Large Language Models (LLMs). Recent research has attempted to improve safety by training models to reason over safety rules before responding. Ho…