paper-with-me

홈 › Papers

Mind the Trojan Horse: Image Prompt Adapter Enabling Scalable and Deceptive Jailbreaking

2025-04-08 · CVPR 2025 1 · Junxi Chen, Junhao Dong, Xiaohua Xie

Recently, the Image Prompt Adapter (IP-Adapter) has been increasingly integrated into text-to-image diffusion models (T2I-DMs) to improve controllability. However, in this paper, we reveal that T2I-DMs equipped with the IP-Adapter (T2I-IP-DMs) enable a new jailbreak attack named the hijacking attack. We demonstrate that, by uploading imperceptible image-space adversarial examples (AEs), the adversary can hijack massive benign users to jailbreak an Image Generation Service (IGS) driven by T2I-IP-DMs and mislead the public to discredit the service provider. Worse still, the IP-Adapter's dependency on open-source image encoders reduces the knowledge required to craft AEs. Extensive experiments verify the technical feasibility of the hijacking attack. In light of the revealed threat, we investigate several existing defenses and explore combining the IP-Adapter with adversarially trained models to overcome existing defenses' limitations. Our code is available at https://github.com/fhdnskfbeuv/attackIPA.

📄 PDF Abstract BibTeX arXiv:2504.05838

Code (1)

fhdnskfbeuv/attackipa 공식 구현 pytorch

Tasks

Image Generation

Methods 이 논문이 사용한 방법론

Golden Queue Managers 설명 없음
Diffusion Diffusion models generate samples by gradually removing noise from a signal, and their training objective can be expressed as a reweighted variational lower-bound…
Adapter 설명 없음

Similar Papers 제목 키워드 기반

Trojan Horse Prompting: Jailbreaking Conversational Multimodal Models by Forging Assistant Message

2025-07-07 · Wei Duan, Li Qian

The rise of conversational interfaces has greatly enhanced LLM usability by leveraging dialogue history for sophisticated reasoning. However, this reliance introduces an unexplored attack surface. This paper introduces T…

Image GenerationSafety Alignment

Trojan horse hunt in deep forecasting models: Insights from the European Space Agency competition

2026-03-20 · Krzysztof Kotowski, Ramez Shendy, Jakub Nalepa, Agata Kaczmarek 외 arxiv

Forecasting plays a crucial role in modern safety-critical applications, such as space operations. However, the increasing use of deep forecasting models introduces a new security risk of trojan horse attacks, carried ou…

Time Series Forecasting

TrojanNet: Exposing the Danger of Trojan Horse Attack on Neural Networks

2020-01-01 · ICLR 2020 1 · Chuan Guo, Ruihan Wu, Kilian Q. Weinberger

The complexity of large-scale neural networks can lead to poor understanding of their internal details. We show that this opaqueness provides an opportunity for adversaries to embed unintended functionalities into the n…

BIG-bench Machine Learning

Trojan Horses in Recruiting: A Red-Teaming Case Study on Indirect Prompt Injection in Standard vs. Reasoning Models

2026-02-19 · Manuel Wirth arxiv

As Large Language Models (LLMs) are increasingly integrated into automated decision-making pipelines, specifically within Human Resources (HR), the security implications of Indirect Prompt Injection (IPI) become critical…

Trojan Horse Training for Breaking Defenses against Backdoor Attacks in Deep Learning

2022-03-25 · Arezoo Rajabi, Bhaskar Ramasubramanian, Radha Poovendran

Machine learning (ML) models that use deep neural networks are vulnerable to backdoor attacks. Such attacks involve the insertion of a (hidden) trigger by an adversary. As a consequence, any input that contains the trigg…

Backdoor Attack