paper-with-me

홈 › Papers

Learning to Detect Unseen Jailbreak Attacks in Large Vision-Language Models

2025-08-08 · Shuang Liang, Zhihao Xu, Jiaqi Weng, Jialing Tao, Hui Xue, Xiting Wang arxiv

Despite extensive alignment efforts, Large Vision-Language Models (LVLMs) remain vulnerable to jailbreak attacks. To mitigate these risks, existing detection methods are essential, yet they face two major challenges: generalization and accuracy. While learning-based methods trained on specific attacks fail to generalize to unseen attacks, learning-free methods based on hand-crafted heuristics suffer from limited accuracy and reduced efficiency. To address these limitations, we propose Learning to Detect (LoD), a learnable framework that eliminates the need for any attack data or hand-crafted heuristics. LoD operates by first extracting layer-wise safety representations directly from the model's internal activations using Multi-modal Safety Concept Activation Vectors classifiers, and then converting the high-dimensional representations into a one-dimensional anomaly score for detection via a Safety Pattern Auto-Encoder. Extensive experiments demonstrate that LoD consistently achieves state-of-the-art detection performance (AUROC) across diverse unseen jailbreak attacks on multiple LVLMs, while also significantly improving efficiency. Code is available at https://github.com/ShuangLiangX/Learning-to-Detect.

📄 PDF Abstract BibTeX arXiv:2508.09201

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Learning to Detect Unknown Jailbreak Attacks in Large Vision-Language Models

2025-10-17 · Shuang Liang, Zhihao Xu, Jialing Tao, Hui Xue 외 arxiv

Despite extensive alignment efforts, Large Vision-Language Models (LVLMs) remain vulnerable to jailbreak attacks, posing serious safety risks. To address this, existing detection methods either learn attack-specific para…

Representation Learning

JailDAM: Jailbreak Detection with Adaptive Memory for Vision-Language Model

2025-04-03 · Yi Nian, Shenzhe Zhu, Yuehan Qin, Li Li 외

Multimodal large language models (MLLMs) excel in vision-language tasks but also pose significant risks of generating harmful content, particularly through jailbreak attacks. Jailbreak attacks refer to intentional manipu…

Language ModelingLanguage Modelling

Beyond Surface-Level Detection: Towards Cognitive-Driven Defense Against Jailbreak Attacks via Meta-Operations Reasoning

2025-08-05 · Rui Pu, Chaozhuo Li, Rui Ha, Litian Zhang 외 arxiv

Defending large language models (LLMs) against jailbreak attacks is essential for their safe and reliable deployment. Existing defenses often rely on shallow pattern matching, which struggles to generalize to novel and u…

Reinforcement Learning

Adversarial Déjà Vu: Jailbreak Dictionary Learning for Stronger Generalization to Unseen Attacks

2025-10-24 · Mahavir Dabas, Tran Huynh, Nikhil Reddy Billa, Jiachen T. Wang 외 arxiv

Large language models remain vulnerable to jailbreak attacks that bypass safety guardrails to elicit harmful outputs. Defending against novel jailbreaks represents a critical challenge in AI safety. Adversarial training …

Adversarial Robustness

Rethinking Jailbreak Detection of Large Vision Language Models with Representational Contrastive Scoring

2025-12-12 · Peichun Hua, Hao Li, Shanghao Shi, Zhiyuan Yu 외 arxiv

Large Vision-Language Models (LVLMs) are vulnerable to a growing array of multimodal jailbreak attacks, necessitating defenses that are both generalizable to novel threats and efficient for practical deployment. Many cur…