paper-with-me

홈 › Papers

Proactive defense against LLM Jailbreak

2025-10-06 · Weiliang Zhao, Jinjun Peng, Daniel Ben-Levi, Zhou Yu, Junfeng Yang arxiv

The proliferation of powerful large language models (LLMs) has necessitated robust safety alignment, yet these models remain vulnerable to evolving adversarial attacks, including multi-turn jailbreaks that iteratively search for successful queries. Current defenses, which are primarily reactive and static, often fail to handle these iterative attacks. In this paper, we introduce ProAct, a novel proactive defense framework designed to disrupt and mislead these iterative search jailbreak methods. Our core idea is to intentionally mislead these jailbreak methods into thinking that the model has been jailbroken with "spurious responses". These misleading responses provide false signals to the attacker's internal optimization loop, causing the adversarial search to terminate prematurely and effectively jailbreaking the jailbreak. By conducting extensive experiments across state-of-the-art LLMs, jailbreaking frameworks, and safety benchmarks, we demonstrate that our method consistently and significantly reduces attack success rates by up to 94% without affecting utility. When combined with other defense fraeworks, it further reduces the latest attack strategies' success rate to 0%. ProActrepresents an orthogonal defense strategy that serves as an additional guardrail to enhance LLM safety against the most effective jailbreaking attacks.

📄 PDF Abstract BibTeX arXiv:2510.05052

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Enhancing Model Defense Against Jailbreaks with Proactive Safety Reasoning

2025-01-31 · Xianglin Yang, Gelei Deng, Jieming Shi, Tianwei Zhang 외

Large language models (LLMs) are vital for a wide range of applications yet remain susceptible to jailbreak threats, which could lead to the generation of inappropriate responses. Conventional defenses, such as refusal a…

BlockingSafety Alignment

Online Learning Defense against Iterative Jailbreak Attacks via Prompt Optimization

2025-10-19 · Masahiro Kaneko, Zeerak Talat, Timothy Baldwin arxiv

Iterative jailbreak methods that repeatedly rewrite and input prompts into large language models (LLMs) to induce harmful outputs -- using the model's previous responses to guide each new iteration -- have been found to …

Reinforcement Learning

HSF: Defending against Jailbreak Attacks with Hidden State Filtering

2024-08-31 · Cheng Qian, Hainan Zhang, Lei Sha, Zhiming Zheng

With the growing deployment of LLMs in daily applications like chatbots and content generation, efforts to ensure outputs align with human values and avoid harmful content have intensified. However, increasingly sophisti…

LLM Jailbreak

Steering Dialogue Dynamics for Robustness against Multi-turn Jailbreaking Attacks

2025-02-28 · Hanjiang Hu, Alexander Robey, Changliu Liu

Large language models (LLMs) are highly vulnerable to jailbreaking attacks, wherein adversarial prompts are designed to elicit harmful responses. While existing defenses effectively mitigate single-turn attacks by detect…

Safety Alignment

Active Honeypot Guardrail System: Probing and Confirming Multi-Turn LLM Jailbreaks

2025-10-16 · ChenYu Wu, Yi Wang, Yang Liao arxiv

Large language models (LLMs) are increasingly vulnerable to multi-turn jailbreak attacks, where adversaries iteratively elicit harmful behaviors that bypass single-turn safety filters. Existing defenses predominantly rel…