paper-with-me

홈 › Papers

Systematic Scaling Analysis of Jailbreak Attacks in Large Language Models

2026-03-11 · Xiangwen Wang, Ananth Balashankar, Varun Chandrasekaran arxiv

Large language models remain vulnerable to jailbreak attacks, yet we still lack a systematic understanding of how jailbreak success scales with attacker effort across methods, model families, and harm types. We initiate a scaling-law framework for jailbreaks by treating each attack as a compute-bounded optimization procedure and measuring progress on a shared FLOPs axis. Our systematic evaluation spans four representative jailbreak paradigms, covering optimization-based attacks, self-refinement prompting, sampling-based selection, and genetic optimization, across multiple model families and scales on a diverse set of harmful goals. We investigate scaling laws that relate attacker budget to attack success score by fitting a simple saturating exponential function to FLOPs--success trajectories, and we derive comparable efficiency summaries from the fitted curves. Empirically, prompting-based paradigms tend to be the most compute-efficient compared to optimization-based methods. To explain this gap, we cast prompt-based updates into an optimization view and show via a same-state comparison that prompt-based attacks more effectively optimize in prompt space. We also show that attacks occupy distinct success--stealthiness operating points with prompting-based methods occupying the high-success, high-stealth region. Finally, we find that vulnerability is strongly goal-dependent: harms involving misinformation are typically easier to elicit than other non-misinformation harms.

📄 PDF Abstract BibTeX arXiv:2603.11149

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

JailbreakEval: An Integrated Toolkit for Evaluating Jailbreak Attempts Against Large Language Models

2024-06-13 · Delong Ran, JinYuan Liu, Yichen Gong, Jingyi Zheng 외

Jailbreak attacks induce Large Language Models (LLMs) to generate harmful responses, posing severe misuse threats. Though research on jailbreak attacks and defenses is emerging, there is no consensus on evaluating jailbr…

How to Trick Your AI TA: A Systematic Study of Academic Jailbreaking in LLM Code Evaluation

2025-12-11 · Devanshu Sahoo, Vasudev Majhi, Arjun Neekhra, Yash Sinha 외 arxiv

The use of Large Language Models (LLMs) as automatic judges for code evaluation is becoming increasingly prevalent in academic environments. But their reliability can be compromised by students who may employ adversarial…

SoK: Robustness in Large Language Models against Jailbreak Attacks

2026-05-06 · Feiyue Xu, Hongsheng Hu, Chaoxiang He, Sheng Hang 외 arxiv

Large Language Models (LLMs) have achieved remarkable success but remain highly susceptible to jailbreak attacks, in which adversarial prompts coerce models into generating harmful, unethical, or policy-violating outputs…

LLM Jailbreak Oracle

2025-06-17 · Shuyi Lin, Anshuman Suri, Alina Oprea, Cheng Tan

As large language models (LLMs) become increasingly deployed in safety-critical applications, the lack of systematic methods to assess their vulnerability to jailbreak attacks presents a critical security gap. We introdu…

LLM Jailbreak

Pruning for Protection: Increasing Jailbreak Resistance in Aligned LLMs Without Fine-Tuning

2024-01-19 · Adib Hasan, Ileana Rugina, Alex Wang

This paper investigates the impact of model compression on the way Large Language Models (LLMs) process prompts, particularly concerning jailbreak resistance. We show that moderate WANDA pruning can enhance resistance to…

Model Compression