paper-with-me

홈 › Papers

Leave My Images Alone: Preventing Multi-Modal Large Language Models from Analyzing Images via Visual Prompt Injection

2026-04-10 · Zedian Shao, Hongbin Liu, Yuepeng Hu, Neil Zhenqiang Gong arxiv

Multi-modal large language models (MLLMs) have emerged as powerful tools for analyzing Internet-scale image data, offering significant benefits but also raising critical safety and societal concerns. In particular, open-weight MLLMs may be misused to extract sensitive information from personal images at scale, such as identities, locations, or other private details. In this work, we propose ImageProtector, a user-side method that proactively protects images before sharing by embedding a carefully crafted, nearly imperceptible perturbation that acts as a visual prompt injection attack on MLLMs. As a result, when an adversary analyzes a protected image with an MLLM, the MLLM is consistently induced to generate a refusal response such as "I'm sorry, I can't help with that request." We empirically demonstrate the effectiveness of ImageProtector across six MLLMs and four datasets. Additionally, we evaluate three potential countermeasures, Gaussian noise, DiffPure, and adversarial training, and show that while they partially mitigate the impact of ImageProtector, they simultaneously degrade model accuracy and/or efficiency. Our study focuses on the practically important setting of open-weight MLLMs and large-scale automated image analysis, and highlights both the promise and the limitations of perturbation-based privacy protection.

📄 PDF Abstract BibTeX arXiv:2604.09024

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

MEDVISTAGYM: A Scalable Training Environment for Thinking with Medical Images via Tool-Integrated Reinforcement Learning

2026-01-12 · Meng Lu, Yuxing Lu, Yuchen Zhuang, Megan Mullins 외 arxiv

Vision language models (VLMs) achieve strong performance on general image understanding but struggle to think with medical images, especially when performing multi-step reasoning through iterative visual interaction. Med…

Reinforcement LearningMultimodal ReasoningVisual Reasoning

Optical Reasoning: Rethinking Images as an Expressive Reasoning Medium Beyond Text

2026-06-08 · Yutong Bian, Dongjie Cheng, Heming Xia, Yongqi Li 외 arxiv

Chain-of-Thought (CoT) improves the performance of Large Language Models (LLMs) and has been extended to Multimodal Large Language Models (MLLMs). More recent work further moves from text-based multimodal reasoning towar…

Multimodal Reasoning

Skywork-R1V4: Toward Agentic Multimodal Intelligence through Interleaved Thinking with Images and DeepResearch

2025-12-02 · Yifan Zhang, Liang Hu, Haofeng Sun, Peiyu Wang 외 arxiv

Despite recent progress in multimodal agentic systems, existing approaches often treat image manipulation and web search as disjoint capabilities, rely heavily on costly reinforcement learning, and lack planning grounded…

Reinforcement LearningImage Manipulation

Bridging Interleaved Multi-Modal Reasoning as a Unified Decision Process

2026-07-04 · Zican Hu, Xuyang Hu, Yiming Liu, Zuwei Long 외 hf

Unified multi-modal models (UMMs) have shown promising interleaved text-image reasoning capabilities, yet effectively optimizing such multi-turn generation via reinforcement learning (RL) remains an open challenge. Exist…

Reinforcement LearningSpatial ReasoningImage GenerationImage Denoising

Mull-Tokens: Modality-Agnostic Latent Thinking

2025-12-11 · Arijit Ray, Ahmed Abdelkader, Chengzhi Mao, Bryan A. Plummer 외 arxiv

Reasoning goes beyond language; the real world requires reasoning about space, time, affordances, and much more that words alone cannot convey. Existing multimodal models exploring the potential of reasoning with images …

Spatial ReasoningVisual Reasoning