paper-with-me

Papers

HarmNet: A Framework for Adaptive Multi-Turn Jailbreak Attacks on Large Language Models

2025-10-21 · Sidhant Narula, Javad Rafiei Asl, Mohammad Ghasemigol, Eduardo Blanco, Daniel Takabi arxiv

Large Language Models (LLMs) remain vulnerable to multi-turn jailbreak attacks. We introduce HarmNet, a modular framework comprising ThoughtNet, a hierarchical semantic network; a feedback-driven Simulator for iterative query refinement; and a Network Traverser for real-time adaptive attack execution. HarmNet systematically explores and refines the adversarial space to uncover stealthy, high-success attack paths. Experiments across closed-source and open-source LLMs show that HarmNet outperforms state-of-the-art methods, achieving higher attack success rates. For example, on Mistral-7B, HarmNet achieves a 99.4% attack success rate, 13.9% higher than the best baseline. Index terms: jailbreak attacks; large language models; adversarial framework; query refinement.

📄 PDF Abstract BibTeX arXiv:2510.18728

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

AutoAdv: Automated Adversarial Prompting for Multi-Turn Jailbreaking of Large Language Models

2025-11-04 · Aashray Reddy, Andrew Zagula, Nicholas Saban arxiv

Large Language Models (LLMs) remain vulnerable to jailbreaking attacks where adversarial prompts elicit harmful outputs. Yet most evaluations focus on single-turn interactions while real-world attacks unfold through adap…

X-Teaming: Multi-Turn Jailbreaks and Defenses with Adaptive Multi-Agents

2025-04-15 · Salman Rahman, Liwei Jiang, James Shiffer, Genglin Liu 외

Multi-turn interactions with language models (LMs) pose critical safety risks, as harmful intent can be strategically spread across exchanges. Yet, the vast majority of prior work has focused on single-turn safety, while…

DiversityRed TeamingSafety Alignment

AJAR: Adaptive Jailbreak Architecture for Red-teaming

2026-01-16 · Yipu Dou, Wang Yang arxiv

Large language model (LLM) safety evaluation is moving from content moderation to action security as modern systems gain persistent state, tool access, and autonomous control loops. Existing jailbreak frameworks still le…

RedTWIZ: Diverse LLM Red Teaming via Adaptive Attack Planning

2025-10-08 · Artur Horal, Daniel Pina, Henrique Paz, Iago Paulo 외 arxiv

This paper presents the vision, scientific contributions, and technical details of RedTWIZ: an adaptive and diverse multi-turn red teaming framework, to audit the robustness of Large Language Models (LLMs) in AI-assisted…

Adversarial AttackRed Teaming

Active Honeypot Guardrail System: Probing and Confirming Multi-Turn LLM Jailbreaks

2025-10-16 · ChenYu Wu, Yi Wang, Yang Liao arxiv

Large language models (LLMs) are increasingly vulnerable to multi-turn jailbreak attacks, where adversaries iteratively elicit harmful behaviors that bypass single-turn safety filters. Existing defenses predominantly rel…