paper-with-me

홈 › Papers

Image-based Prompt Injection: Hijacking Multimodal LLMs through Visually Embedded Adversarial Instructions

2026-03-04 · Neha Nagaraja, Lan Zhang, Zhilong Wang, Bo Zhang, Pawan Patil arxiv

Multimodal Large Language Models (MLLMs) integrate vision and text to power applications, but this integration introduces new vulnerabilities. We study Image-based Prompt Injection (IPI), a black-box attack in which adversarial instructions are embedded into natural images to override model behavior. Our end-to-end IPI pipeline incorporates segmentation-based region selection, adaptive font scaling, and background-aware rendering to conceal prompts from human perception while preserving model interpretability. Using the COCO dataset and GPT-4-turbo, we evaluate 12 adversarial prompt strategies and multiple embedding configurations. The results show that IPI can reliably manipulate the output of the model, with the most effective configuration achieving up to 64\% attack success under stealth constraints. These findings highlight IPI as a practical threat in black-box settings and underscore the need for defenses against multimodal prompt injection.

📄 PDF Abstract BibTeX arXiv:2603.03637

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Pseudo-Conversation Injection for LLM Goal Hijacking

2024-10-31 · Zheng Chen, Buhui Yao

Goal hijacking is a type of adversarial attack on Large Language Models (LLMs) where the objective is to manipulate the model into producing a specific, predetermined output, regardless of the user's original input. In g…

Adversarial Attack

Beyond the Benchmark: Innovative Defenses Against Prompt Injection Attacks

2025-12-18 · Safwan Shaheer, G. M. Refatul Islam, Mohammad Rafid Hamid, Tahsin Zaman Jilan arxiv

In this fast-evolving area of LLMs, our paper discusses the significant security risk presented by prompt injection attacks. It focuses on small open-sourced models, specifically the LLaMA family of models. We introduce …

Empirical Analysis of Large Vision-Language Models against Goal Hijacking via Visual Prompt Injection

2024-08-07 · Subaru Kimura, Ryota Tanaka, Shumpei Miyawaki, Jun Suzuki 외

We explore visual prompt injection (VPI) that maliciously exploits the ability of large vision-language models (LVLMs) to follow instructions drawn onto the input image. We propose a new VPI method, "goal hijacking via v…

Instruction Following

Tensor Trust: Interpretable Prompt Injection Attacks from an Online Game

2023-11-02 · Sam Toyer, Olivia Watkins, Ethan Adrian Mendes, Justin Svegliato 외

While Large Language Models (LLMs) are increasingly being used in real-world applications, they remain vulnerable to prompt injection attacks: malicious third party prompts that subvert the intent of the system designer.…

Instruction Following

CHAI: Command Hijacking against embodied AI

2025-09-30 · Luis Burbano, Diego Ortiz, Qi Sun, Siwei Yang 외 arxiv

Embodied Artificial Intelligence (AI) promises to handle edge cases in robotic vehicle systems where data is scarce by using common-sense reasoning grounded in perception and action to generalize beyond training distribu…

Adversarial RobustnessMultimodal ReasoningAutonomous DrivingObject Tracking