paper-with-me

홈 › Papers

F2A: An Innovative Approach for Prompt Injection by Utilizing Feign Security Detection Agents

2024-10-11 · Yupeng Ren

With the rapid development of Large Language Models (LLMs), numerous mature applications of LLMs have emerged in the field of content safety detection. However, we have found that LLMs exhibit blind trust in safety detection agents. The general LLMs can be compromised by hackers with this vulnerability. Hence, this paper proposed an attack named Feign Agent Attack (F2A).Through such malicious forgery methods, adding fake safety detection results into the prompt, the defense mechanism of LLMs can be bypassed, thereby obtaining harmful content and hijacking the normal conversation. Continually, a series of experiments were conducted. In these experiments, the hijacking capability of F2A on LLMs was analyzed and demonstrated, exploring the fundamental reasons why LLMs blindly trust safety detection results. The experiments involved various scenarios where fake safety detection results were injected into prompts, and the responses were closely monitored to understand the extent of the vulnerability. Also, this paper provided a reasonable solution to this attack, emphasizing that it is important for LLMs to critically evaluate the results of augmented agents to prevent the generating harmful content. By doing so, the reliability and security can be significantly improved, protecting the LLMs from F2A.

📄 PDF Abstract BibTeX arXiv:2410.08776

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Beyond the Benchmark: Innovative Defenses Against Prompt Injection Attacks

2025-12-18 · Safwan Shaheer, G. M. Refatul Islam, Mohammad Rafid Hamid, Tahsin Zaman Jilan arxiv

In this fast-evolving area of LLMs, our paper discusses the significant security risk presented by prompt injection attacks. It focuses on small open-sourced models, specifically the LLaMA family of models. We introduce …

Assessing Prompt Injection Risks in 200+ Custom GPTs

2023-11-20 · Jiahao Yu, Yuhang Wu, Dong Shu, Mingyu Jin 외

In the rapidly evolving landscape of artificial intelligence, ChatGPT has been widely used in various applications. The new feature - customization of ChatGPT models by users to cater to specific needs has opened new fro…

Prompt Injection 2.0: Hybrid AI Threats

2025-07-17 · Jeremy McHugh, Kristina Šekrst, Jon Cefalu

Prompt injection attacks, where malicious input is designed to manipulate AI systems into ignoring their original instructions and following unauthorized commands instead, were first discovered by Preamble, Inc. in May 2…

Trust No AI: Prompt Injection Along The CIA Security Triad

2024-12-08 · Johann Rehberger

The CIA security triad - Confidentiality, Integrity, and Availability - is a cornerstone of data and cybersecurity. With the emergence of large language model (LLM) applications, a new class of threat, known as prompt in…

Language ModelingLanguage ModellingLarge Language Model

Meta SecAlign: A Secure Foundation LLM Against Prompt Injection Attacks

2025-07-03 · Sizhe Chen, Arman Zharmagambetov, David Wagner, Chuan Guo

Prompt injection attacks pose a significant security threat to LLM-integrated applications. Model-level defenses have shown strong effectiveness, but are currently deployed into commercial-grade models in a closed-source…

Instruction Following