paper-with-me

Papers

Beyond Visual Safety: Jailbreaking Multimodal Large Language Models for Harmful Image Generation via Semantic-Agnostic Inputs

2026-01-22 · Mingyu Yu, Lana Liu, Zhehao Zhao, Wei Wang, Sujuan Qin arxiv

The rapid advancement of Multimodal Large Language Models (MLLMs) has introduced complex security challenges, particularly at the intersection of textual and visual safety. While existing schemes have explored the security vulnerabilities of MLLMs, the investigation into their visual safety boundaries remains insufficient. In this paper, we propose Beyond Visual Safety (BVS), a novel image-text pair jailbreaking framework specifically designed to probe the visual safety boundaries of MLLMs. BVS employs a "reconstruction-then-generation" strategy, leveraging neutralized visual splicing and inductive recomposition to decouple malicious intent from raw inputs, thereby leading MLLMs to be induced into generating harmful images. Experimental results demonstrate that BVS achieves a remarkable jailbreak success rate of 98.21\% against GPT-5 (12 January 2026 release). Our findings expose critical vulnerabilities in the visual safety alignment of current MLLMs.

📄 PDF Abstract BibTeX arXiv:2601.15698

Code (0)

등록된 구현이 없습니다.

Tasks

Image Generation

Similar Papers 제목 키워드 기반

KSAFE-MM: A Multimodal Safety Benchmark via Localized Contextualization for Korean Cultural Risks

2026-05-27 · Yongwoo Kim, Sojung An, Yunjin Park, Jungwon Yoon 외 arxiv

Multimodal Large Language Models (MLLMs) exacerbate safety risks by introducing vulnerabilities across multiple modalities, such as language and vision. Current MLLM safety evaluation tools, however, suffer from major li…

DMN: A Compositional Framework for Jailbreaking Multimodal LLMs with Multi-Image Inputs

2026-05-18 · Wenzhuo Xu, Zhipeng Wei, Zonghao Ying, Deyue Zhang 외 arxiv

Multimodal Large Language Models (MLLMs) are vulnerable to jailbreak attacks, which can elicit harmful responses from MLLMs. Many MLLMs support multi-image inputs, inadvertently introducing new vulnerabilities due to les…

Visual Reasoning

Sequential Comics for Jailbreaking Multimodal Large Language Models via Structured Visual Storytelling

2025-10-16 · Deyue Zhang, Dongdong Yang, Junjie Mu, Quancheng Zou 외 arxiv

Multimodal large language models (MLLMs) exhibit remarkable capabilities but remain susceptible to jailbreak attacks exploiting cross-modal vulnerabilities. In this work, we introduce a novel method that leverages sequen…

Visual Storytelling

Preventing Robotic Jailbreaking via Multimodal Domain Adaptation

2025-09-27 · Francesco Marchiori, Rohan Sinha, Christopher Agia, Alexander Robey 외 arxiv

Large Language Models (LLMs) and Vision-Language Models (VLMs) are increasingly deployed in robotic environments but remain vulnerable to jailbreaking attacks that bypass safety mechanisms and drive unsafe or physically …

Autonomous DrivingDomain Adaptation

Distraction is All You Need for Multimodal Large Language Model Jailbreaking

2025-02-15 · CVPR 2025 1 · Zuopeng Yang, Jiluan Fan, Anli Yan, Erdun Gao 외

Multimodal Large Language Models (MLLMs) bridge the gap between visual and textual data, enabling a range of advanced applications. However, complex internal interactions among visual elements and their alignment with te…

AllLanguage ModelingLanguage ModellingLarge Language Model+1