paper-with-me

홈 › Papers

Learning to Detect Unknown Jailbreak Attacks in Large Vision-Language Models

2025-10-17 · Shuang Liang, Zhihao Xu, Jialing Tao, Hui Xue, Xiting Wang arxiv

Despite extensive alignment efforts, Large Vision-Language Models (LVLMs) remain vulnerable to jailbreak attacks, posing serious safety risks. To address this, existing detection methods either learn attack-specific parameters, which hinders generalization to unseen attacks, or rely on heuristically sound principles, which limit accuracy and efficiency. To overcome these limitations, we propose Learning to Detect (LoD), a general framework that accurately detects unknown jailbreak attacks by shifting the focus from attack-specific learning to task-specific learning. This framework includes a Multi-modal Safety Concept Activation Vector module for safety-oriented representation learning and a Safety Pattern Auto-Encoder module for unsupervised attack classification. Extensive experiments show that our method achieves consistently higher detection AUROC on diverse unknown attacks while improving efficiency. The code is available at https://anonymous.4open.science/r/Learning-to-Detect-51CB.

📄 PDF Abstract BibTeX arXiv:2510.15430

Code (0)

등록된 구현이 없습니다.

Tasks

Representation Learning

Similar Papers 제목 키워드 기반

Learning to Detect Unseen Jailbreak Attacks in Large Vision-Language Models

2025-08-08 · Shuang Liang, Zhihao Xu, Jiaqi Weng, Jialing Tao 외 arxiv

Despite extensive alignment efforts, Large Vision-Language Models (LVLMs) remain vulnerable to jailbreak attacks. To mitigate these risks, existing detection methods are essential, yet they face two major challenges: gen…

JailDAM: Jailbreak Detection with Adaptive Memory for Vision-Language Model

2025-04-03 · Yi Nian, Shenzhe Zhu, Yuehan Qin, Li Li 외

Multimodal large language models (MLLMs) excel in vision-language tasks but also pose significant risks of generating harmful content, particularly through jailbreak attacks. Jailbreak attacks refer to intentional manipu…

Language ModelingLanguage Modelling

HiddenDetect: Detecting Jailbreak Attacks against Large Vision-Language Models via Monitoring Hidden States

2025-02-20 · Yilei Jiang, Xinyan Gao, Tianshuo Peng, Yingshui Tan 외

The integration of additional modalities increases the susceptibility of large vision-language models (LVLMs) to safety risks, such as jailbreak attacks, compared to their language-only counterparts. While existing resea…

Robust Prompt Optimization for Defending Language Models Against Jailbreaking Attacks

2024-01-30 · Andy Zhou, Bo Li, Haohan Wang

Despite advances in AI alignment, large language models (LLMs) remain vulnerable to adversarial attacks or jailbreaking, in which adversaries can modify prompts to induce unwanted behavior. While some defenses have been …

The Great Contradiction Showdown: How Jailbreak and Stealth Wrestle in Vision-Language Models?

2024-10-02 · Ching-Chia Kao, Chia-Mu Yu, Chun-Shien Lu, Chu-Song Chen

Vision-Language Models (VLMs) have achieved remarkable performance on a variety of tasks, yet they remain vulnerable to jailbreak attacks that compromise safety and reliability. In this paper, we provide an information-t…