paper-with-me

홈 › Papers

The VLLM Safety Paradox: Dual Ease in Jailbreak Attack and Defense

2024-11-13 · Yangyang Guo, Fangkai Jiao, Liqiang Nie, Mohan Kankanhalli

The vulnerability of Vision Large Language Models (VLLMs) to jailbreak attacks appears as no surprise. However, recent defense mechanisms against these attacks have reached near-saturation performance on benchmark evaluations, often with minimal effort. This \emph{dual high performance} in both attack and defense raises a fundamental and perplexing paradox. To gain a deep understanding of this issue and thus further help strengthen the trustworthiness of VLLMs, this paper makes three key contributions: i) One tentative explanation for VLLMs being prone to jailbreak attacks--\textbf{inclusion of vision inputs}, as well as its in-depth analysis. ii) The recognition of a largely ignored problem in existing defense mechanisms--\textbf{over-prudence}. The problem causes these defense methods to exhibit unintended abstention, even in the presence of benign inputs, thereby undermining their reliability in faithfully defending against attacks. iii) A simple safety-aware method--\textbf{LLM-Pipeline}. Our method repurposes the more advanced guardrails of LLMs on the shelf, serving as an effective alternative detector prior to VLLM response. Last but not least, we find that the two representative evaluation methods for jailbreak often exhibit chance agreement. This limitation makes it potentially misleading when evaluating attack strategies or defense mechanisms. We believe the findings from this paper offer useful insights to rethink the foundational development of VLLM safety with respect to benchmark datasets, defense strategies, and evaluation methods.

📄 PDF Abstract BibTeX arXiv:2411.08410

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

How Many Unicorns Are in This Image? A Safety Evaluation Benchmark for Vision LLMs

2023-11-27 · Haoqin Tu, Chenhang Cui, Zijun Wang, Yiyang Zhou 외

This work focuses on the potential of Vision LLMs (VLLMs) in visual reasoning. Different from prior studies, we shift our focus from evaluating standard performance to introducing a comprehensive safety evaluation suite,…

Adversarial RobustnessVisual Question Answering (VQA)Visual Reasoning

Safety Fine-Tuning at (Almost) No Cost: A Baseline for Vision Large Language Models

2024-02-03 · Yongshuo Zong, Ondrej Bohdal, Tingyang Yu, Yongxin Yang 외

Current vision large language models (VLLMs) exhibit remarkable capabilities yet are prone to generate harmful content and are vulnerable to even the simplest jailbreaking attacks. Our initial analysis finds that this is…

Instruction FollowingSafety Alignment

[WIP] Jailbreak Paradox: The Achilles' Heel of LLMs

2024-06-18 · Abhinav Rao, Monojit Choudhury, Somak Aditya

We introduce two paradoxes concerning jailbreak of foundation models: First, it is impossible to construct a perfect jailbreak classifier, and second, a weaker model cannot consistently detect whether a stronger (in a pa…

Safety Paradox: How Enhanced Safety Awareness Leaves LLMs Vulnerable to Posterior Attack

2026-06-04 · Long P. Hoang, Hai V. Le, Shaoyang Xu, Wei Lu 외 arxiv

Large language models (LLMs) are rigorously aligned to refuse harmful requests, a process that inherently cultivates a latent capacity to evaluate and recognize unsafe content. In this work, we reveal that this advanced …

Reinforcement Learning

Are Vision LLMs Road-Ready? A Comprehensive Benchmark for Safety-Critical Driving Video Understanding

2025-04-20 · Tong Zeng, Longfeng Wu, Liang Shi, Dawei Zhou 외

Vision Large Language Models (VLLMs) have demonstrated impressive capabilities in general visual tasks such as image captioning and visual question answering. However, their effectiveness in specialized, safety-critical …

Autonomous DrivingImage CaptioningMultiple-choiceQuestion Answering+3