paper-with-me

홈 › Papers

Trojan Horse Prompting: Jailbreaking Conversational Multimodal Models by Forging Assistant Message

2025-07-07 · Wei Duan, Li Qian

The rise of conversational interfaces has greatly enhanced LLM usability by leveraging dialogue history for sophisticated reasoning. However, this reliance introduces an unexplored attack surface. This paper introduces Trojan Horse Prompting, a novel jailbreak technique. Adversaries bypass safety mechanisms by forging the model's own past utterances within the conversational history provided to its API. A malicious payload is injected into a model-attributed message, followed by a benign user prompt to trigger harmful content generation. This vulnerability stems from Asymmetric Safety Alignment: models are extensively trained to refuse harmful user requests but lack comparable skepticism towards their own purported conversational history. This implicit trust in its "past" creates a high-impact vulnerability. Experimental validation on Google's Gemini-2.0-flash-preview-image-generation shows Trojan Horse Prompting achieves a significantly higher Attack Success Rate (ASR) than established user-turn jailbreaking methods. These findings reveal a fundamental flaw in modern conversational AI security, necessitating a paradigm shift from input-level filtering to robust, protocol-level validation of conversational context integrity.

📄 PDF Abstract BibTeX arXiv:2507.04673

Code (0)

등록된 구현이 없습니다.

Tasks

Image GenerationSafety Alignment

Similar Papers 제목 키워드 기반

Jailbreaking Leading Safety-Aligned LLMs with Simple Adaptive Attacks

2024-04-02 · Maksym Andriushchenko, Francesco Croce, Nicolas Flammarion

We show that even the most recent safety-aligned LLMs are not robust to simple adaptive jailbreaking attacks. First, we demonstrate how to successfully leverage access to logprobs for jailbreaking: we initially design an…

In-Context Learning

Trojan horse hunt in deep forecasting models: Insights from the European Space Agency competition

2026-03-20 · Krzysztof Kotowski, Ramez Shendy, Jakub Nalepa, Agata Kaczmarek 외 arxiv

Forecasting plays a crucial role in modern safety-critical applications, such as space operations. However, the increasing use of deep forecasting models introduces a new security risk of trojan horse attacks, carried ou…

Time Series Forecasting

TrojanNet: Exposing the Danger of Trojan Horse Attack on Neural Networks

2020-01-01 · ICLR 2020 1 · Chuan Guo, Ruihan Wu, Kilian Q. Weinberger

The complexity of large-scale neural networks can lead to poor understanding of their internal details. We show that this opaqueness provides an opportunity for adversaries to embed unintended functionalities into the n…

BIG-bench Machine Learning

The AI Cognitive Trojan Horse: How Large Language Models May Bypass Human Epistemic Vigilance

2026-01-11 · Andrew D. Maynard arxiv

Large language model (LLM)-based conversational AI systems present a challenge to human cognition that current frameworks for understanding misinformation and persuasion do not adequately address. This paper proposes tha…

Mind the Trojan Horse: Image Prompt Adapter Enabling Scalable and Deceptive Jailbreaking

2025-04-08 · CVPR 2025 1 · Junxi Chen, Junhao Dong, Xiaohua Xie

Recently, the Image Prompt Adapter (IP-Adapter) has been increasingly integrated into text-to-image diffusion models (T2I-DMs) to improve controllability. However, in this paper, we reveal that T2I-DMs equipped with the …

Image Generation