paper-with-me

홈 › Papers

An Empirical Study of Multi-Generation Sampling for Jailbreak Detection in Large Language Models

2026-04-20 · Hanrui Luo, Shreyank N Gowda arxiv

Detecting jailbreak behaviour in large language models remains challenging, particularly when strongly aligned models produce harmful outputs only rarely. In this work, we present an empirical study of output based jailbreak detection under realistic conditions using the JailbreakBench Behaviors dataset and multiple generator models with varying alignment strengths. We evaluate both a lexical TF-IDF detector and a generation inconsistency based detector across different sampling budgets. Our results show that single output evaluation systematically underestimates jailbreak vulnerability, as increasing the number of sampled generations reveals additional harmful behaviour. The most significant improvements occur when moving from a single generation to moderate sampling, while larger sampling budgets yield diminishing returns. Cross generator experiments demonstrate that detection signals partially generalise across models, with stronger transfer observed within related model families. A category level analysis further reveals that lexical detectors capture a mixture of behavioural signals and topic specific cues, rather than purely harmful behaviour. Overall, our findings suggest that moderate multi sample auditing provides a more reliable and practical approach for estimating model vulnerability and improving jailbreak detection in large language models. Code will be released.

📄 PDF Abstract BibTeX arXiv:2604.18775

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

A Cross-Language Investigation into Jailbreak Attacks in Large Language Models

2024-01-30 · Jie Li, Yi Liu, Chongyang Liu, Ling Shi 외

Large Language Models (LLMs) have become increasingly popular for their advanced text generation capabilities across various domains. However, like any software, they face security challenges, including the risk of 'jail…

Text Generation

Step-Wise Refusal Dynamics in Autoregressive and Diffusion Language Models

2026-02-01 · Eliron Rahimi, Elad Hirshel, Rom Himelstein, Amit LeVi 외 arxiv

Diffusion language models (DLMs) have recently emerged as a competitive alternative to autoregressive (AR) models, offering parallel decoding, competitive generation quality, and initial evidence of improved jailbreak ro…

Jailbreaking ChatGPT via Prompt Engineering: An Empirical Study

2023-05-23 · Yi Liu, Gelei Deng, Zhengzi Xu, Yuekang Li 외

Large Language Models (LLMs), like ChatGPT, have demonstrated vast potential but also introduce challenges related to content constraints and potential misuse. Our study investigates three key research questions: (1) the…

Prompt Engineering

Don't Listen To Me: Understanding and Exploring Jailbreak Prompts of Large Language Models

2024-03-26 · Zhiyuan Yu, Xiaogeng Liu, Shunning Liang, Zach Cameron 외

Recent advancements in generative AI have enabled ubiquitous access to large language models (LLMs). Empowered by their exceptional capabilities to understand and generate human-like text, these models are being increasi…

Multi-Turn Jailbreaks Are Simpler Than They Seem

2025-08-11 · Xiaoxue Yang, Jaeha Lee, Anna-Katharina Dick, Jasper Timm 외 arxiv

While defenses against single-turn jailbreak attacks on Large Language Models (LLMs) have improved significantly, multi-turn jailbreaks remain a persistent vulnerability, often achieving success rates exceeding 70% again…