paper-with-me

홈 › Papers

Reasoned Safety Alignment: Ensuring Jailbreak Defense via Answer-Then-Check

2025-09-15 · Chentao Cao, Xiaojun Xu, Bo Han, Hang Li arxiv

As large language models (LLMs) continue to advance in capabilities, ensuring their safety against jailbreak attacks remains a critical challenge. In this paper, we introduce a novel safety alignment approach called Answer-Then-Check, which enhances LLM robustness against malicious prompts by applying thinking ability to mitigate jailbreaking problems before producing a final answer to the user. Our method enables models to answer the question in their thoughts directly and then critically evaluate its safety before deciding whether to provide it. To implement this approach, we construct the Reasoned Safety Alignment (ReSA) dataset, comprising 80K samples that teach models to reason through direct responses and then analyze their safety. Experimental results demonstrate that our approach achieves the Pareto frontier with superior safety capability while decreasing over-refusal rates. Notably, the fine-tuned model maintains general reasoning capabilities on benchmarks like MMLU, MATH500, and HumanEval. Besides, our method equips models with the ability to perform safe completion, while post-hoc detection methods can only directly reject sensitive, harmful queries (e.g., self-harm). Our results show that inference-time strategies alone are insufficient, highlighting the necessity of safety training, and we find even $500$ samples can yield performance comparable to the entire dataset, suggesting a promising path for data-efficient safety alignment. The dataset is publicly available at: https://huggingface.co/datasets/ByteDance-Seed/ReSA.

📄 PDF Abstract BibTeX arXiv:2509.11629

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Steering Dialogue Dynamics for Robustness against Multi-turn Jailbreaking Attacks

2025-02-28 · Hanjiang Hu, Alexander Robey, Changliu Liu

Large language models (LLMs) are highly vulnerable to jailbreaking attacks, wherein adversarial prompts are designed to elicit harmful responses. While existing defenses effectively mitigate single-turn attacks by detect…

Safety Alignment

MUSE: MCTS-Driven Red Teaming Framework for Enhanced Multi-Turn Dialogue Safety in Large Language Models

2025-09-18 · Siyu Yan, Long Zeng, Xuecheng Wu, Chengcheng Han 외 arxiv

As large language models~(LLMs) become widely adopted, ensuring their alignment with human values is crucial to prevent jailbreaks where adversaries manipulate models to produce harmful content. While most defenses targe…

Red Teaming

Unified Defense for Large Language Models against Jailbreak and Fine-Tuning Attacks in Education

2025-11-18 · Xin Yi, Yue Li, Dongsheng Shi, Linlin Wang 외 arxiv

Large Language Models (LLMs) are increasingly integrated into educational applications. However, they remain vulnerable to jailbreak and fine-tuning attacks, which can compromise safety alignment and lead to harmful outp…

SafeAligner: Safety Alignment against Jailbreak Attacks via Response Disparity Guidance

2024-06-26 · Caishuang Huang, Wanxu Zhao, Rui Zheng, Huijie Lv 외

As the development of large language models (LLMs) rapidly advances, securing these models effectively without compromising their utility has become a pivotal area of research. However, current defense strategies against…

Safety Alignment

Latent-space adversarial training with post-aware calibration for defending large language models against jailbreak attacks

2025-01-18 · Xin Yi, Yue Li, Dongsheng Shi, LinLin Wang 외

Ensuring safety alignment is a critical requirement for large language models (LLMs), particularly given increasing deployment in real-world applications. Despite considerable advancements, LLMs remain susceptible to jai…

Safety Alignment