paper-with-me

Papers

Jailbreaking Multimodal Large Language Models using Multi-Clip Video

2026-06-01 · Choongwon Kang, Seungjong Sun, Hyunmin Jun, Jang Hyun Kim arxiv

As multimodal large language models (MLLMs) have advanced to process video inputs, concerns have emerged about their potential for malicious misuse. Prior jailbreak studies have shown that safety alignment in MLLMs can be bypassed through visual inputs, yet it remains unclear which properties of video inputs induce this vulnerability. To address this gap, we introduce Multi-Clip Video (MCV) SafetyBench, a dataset of 2,920 videos designed to evaluate how the diversity of video inputs affects the vulnerability of MLLMs. Each video consists of multiple short clips depicting diverse contexts related to a harmful query. Experiments on eight representative video MLLMs show that attack success consistently increases with the number of clips. Our results further indicate that the video modality is (1) more vulnerable than the image modality, (2) more vulnerable to dynamic videos than to static videos, and (3) more vulnerable when videos contain more diverse contexts. Building on these findings, we propose a defense strategy that leverages the relative robustness of the image modality.

📄 PDF Abstract BibTeX arXiv:2606.02111

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

From LLMs to MLLMs: Exploring the Landscape of Multimodal Jailbreaking

2024-06-21 · Siyuan Wang, Zhuohan Long, Zhihao Fan, Zhongyu Wei

The rapid development of Large Language Models (LLMs) and Multimodal Large Language Models (MLLMs) has exposed vulnerabilities to various adversarial attacks. This paper provides a comprehensive overview of jailbreaking …

Divide and Conquer: A Hybrid Strategy Defeats Multimodal Large Language Models

2024-12-21 · Yanxu Mao, Peipei Liu, Tiehan Cui, Congying Liu 외

Large language models (LLMs) are widely applied in various fields of society due to their powerful reasoning, understanding, and generation capabilities. However, the security issues associated with these models are beco…

Efficient Indirect LLM Jailbreak via Multimodal-LLM Jailbreak

2024-05-30 · Zhenxing Niu, Yuyao Sun, Haoxuan Ji, Zheng Lin 외

This paper focuses on jailbreaking attacks against large language models (LLMs), eliciting them to generate objectionable content in response to harmful user queries. Unlike previous LLM-jailbreak methods that directly o…

Language ModelingLanguage ModellingLarge Language ModelLLM Jailbreak+1

Continuous Embedding Attacks via Clipped Inputs in Jailbreaking Large Language Models

2024-07-16 · Zihao Xu, Yi Liu, Gelei Deng, Kailong Wang 외

Security concerns for large language models (LLMs) have recently escalated, focusing on thwarting jailbreaking attempts in discrete prompts. However, the exploration of jailbreak vulnerabilities arising from continuous e…

Jailbreaking Large Vision Language Models in Intelligent Transportation Systems

2025-11-17 · Badhan Chandra Das, Md Tasnim Jawad, Md Jueal Mia, M. Hadi Amini 외 arxiv

Large Vision Language Models (LVLMs) demonstrate strong capabilities in multimodal reasoning and many real-world applications, such as visual question answering. However, LVLMs are highly vulnerable to jailbreaking attac…

Visual Question AnsweringMultimodal Reasoning