paper-with-me

홈 › Papers

The Great Contradiction Showdown: How Jailbreak and Stealth Wrestle in Vision-Language Models?

2024-10-02 · Ching-Chia Kao, Chia-Mu Yu, Chun-Shien Lu, Chu-Song Chen

Vision-Language Models (VLMs) have achieved remarkable performance on a variety of tasks, yet they remain vulnerable to jailbreak attacks that compromise safety and reliability. In this paper, we provide an information-theoretic framework for understanding the fundamental trade-off between the effectiveness of these attacks and their stealthiness. Drawing on Fano's inequality, we demonstrate how an attacker's success probability is intrinsically linked to the stealthiness of generated prompts. Building on this, we propose an efficient algorithm for detecting non-stealthy jailbreak attacks, offering significant improvements in model robustness. Experimental results highlight the tension between strong attacks and their detectability, providing insights into both adversarial strategies and defense mechanisms.

📄 PDF Abstract BibTeX arXiv:2410.01438

Code (0)

등록된 구현이 없습니다.

Methods 이 논문이 사용한 방법론

Diffusion Diffusion models generate samples by gradually removing noise from a signal, and their training objective can be expressed as a reweighted variational lower-bound…

Similar Papers 제목 키워드 기반

Bad Weather, Social Network, and Internal Migration; Case of Japanese Sumo Wrestlers 1946-1985

2022-04-17 · Eiji Yamamura

Post-World War II , there was massive internal migration from rural to urban areas in Japan. The location of Sumo stables was concentrated in Tokyo. Hence, supply of Sumo wrestlers from rural areas to Tokyo was considere…

When Safety Detectors Aren't Enough: A Stealthy and Effective Jailbreak Attack on LLMs via Steganographic Techniques

2025-05-22 · Jianing Geng, Biao Yi, Zekun Fei, Tongxi Wu 외

Jailbreak attacks pose a serious threat to large language models (LLMs) by bypassing built-in safety mechanisms and leading to harmful outputs. Studying these attacks is crucial for identifying vulnerabilities and improv…

Benchmarking

AutoDAN: Generating Stealthy Jailbreak Prompts on Aligned Large Language Models

2023-10-03 · Xiaogeng Liu, Nan Xu, Muhao Chen, Chaowei Xiao

The aligned Large Language Models (LLMs) are powerful language understanding and decision-making tools that are created through extensive alignment with human feedback. However, these large models remain susceptible to j…

Decision Making

Injecting Universal Jailbreak Backdoors into LLMs in Minutes

2025-02-09 · Zhuowei Chen, Qiannan Zhang, Shichao Pei

Jailbreak backdoor attacks on LLMs have garnered attention for their effectiveness and stealth. However, existing methods rely on the crafting of poisoned datasets and the time-consuming process of fine-tuning. In this w…

Model Editing

AudioJailbreak: Jailbreak Attacks against End-to-End Large Audio-Language Models

2025-05-20 · Guangke Chen, Fu Song, Zhe Zhao, Xiaojun Jia 외

Jailbreak attacks to Large audio-language models (LALMs) are studied recently, but they achieve suboptimal effectiveness, applicability, and practicability, particularly, assuming that the adversary can fully manipulate …

text-to-speechText to Speech