paper-with-me

홈 › Papers

Hard to Read, Easy to Jailbreak: How Visual Degradation Bypasses MLLM Safety Alignment

2026-05-08 · Zhixue Song, Boyan Han, Yiwei Wang, Chi Zhang arxiv

Recent advancements in visual context compression enable MLLMs to process ultra-long contexts efficiently by rendering text into images. However, we identify a critical vulnerability inherent to this paradigm: lowering image resolution inadvertently catalyzes jailbreaking. Our experiments reveal that the safety defenses of SOTA models deteriorate sharply as resolution degrades, surprisingly persisting even when text remains legible. We attribute this to `Cognitive Overload'', hypothesizing that the effort required to decipher degraded inputs diverts attentional resources from safety auditing. This phenomenon is consistent across various visual perturbations, including noise and geometric distortion. To address this, we propose a simple `Structured Cognitive Offloading'' strategy that mitigates these risks by enforcing a serialized pipeline to decouple visual transcription from safety assessment. Our work exposes a significant risk in vision-based compression and provides critical insights for the secure design of future MLLMs.

📄 PDF Abstract BibTeX arXiv:2605.07250

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Improved Techniques for Optimization-Based Jailbreaking on Large Language Models

2024-05-31 · Xiaojun Jia, Tianyu Pang, Chao Du, Yihao Huang 외

Large language models (LLMs) are being rapidly developed, and a key component of their widespread deployment is their safety-related alignment. Many red-teaming efforts aim to jailbreak LLMs, where among these efforts, t…

Red Teaming

The Jailbreak Tax: How Useful are Your Jailbreak Outputs?

2025-04-14 · Kristina Nikolić, Luze Sun, Jie Zhang, Florian Tramèr

Jailbreak attacks bypass the guardrails of large language models to produce harmful outputs. In this paper, we ask whether the model outputs produced by existing jailbreaks are actually useful. For example, when jailbrea…

Math

AutoDAN: Interpretable Gradient-Based Adversarial Attacks on Large Language Models

2023-10-23 · Sicheng Zhu, Ruiyi Zhang, Bang An, Gang Wu 외

Safety alignment of Large Language Models (LLMs) can be compromised with manual jailbreak attacks and (automatic) adversarial attacks. Recent studies suggest that defending against these attacks is possible: adversarial …

Adversarial AttackBlockingSafety Alignment

EasyJailbreak: A Unified Framework for Jailbreaking Large Language Models

2024-03-18 · Weikang Zhou, Xiao Wang, Limao Xiong, Han Xia 외

Jailbreak attacks are crucial for identifying and mitigating the security vulnerabilities of Large Language Models (LLMs). They are designed to bypass safeguards and elicit prohibited outputs. However, due to significant…

Test-Time Immunization: A Universal Defense Framework Against Jailbreaks for (Multimodal) Large Language Models

2025-05-28 · Yongcan Yu, Yanbo Wang, Ran He, Jian Liang

While (multimodal) large language models (LLMs) have attracted widespread attention due to their exceptional capabilities, they remain vulnerable to jailbreak attacks. Various defense methods are proposed to defend again…