paper-with-me

홈 › Papers

Evaluation of Prompt Injection Defenses in Large Language Models

2026-04-26 · Priyal Deep, Shane Emmons, Amy Fox, Kyle Bacon, Kelley McAllister, Peter Ortiz, Krisztian Flautner arxiv

LLM-powered applications routinely embed secrets in system prompts, yet models can be tricked into revealing them. We built an adaptive attacker that evolves its strategies over hundreds of rounds and tested it against nine defense configurations across more than 20,000 attacks. Every defense that relied on the model to protect itself eventually broke. The only defense that held was output filtering, which checks the model's responses via hardcoded rules in separate application code before they reach the user, achieving zero leaks across 15,000 attacks. These results demonstrate that security boundaries must be enforced in application code, not by the model being attacked. Until such defenses are verified by tools like Swept AI, AI systems handling sensitive operations should be restricted to internal, trusted personnel.

📄 PDF Abstract BibTeX arXiv:2604.23887

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

LongPIBench: A Long-Context Benchmark for Prompt Injection

2026-08-28 · Yupei Liu, Yuqi Jia, Neil Zhenqiang Gong, Jinyuan Jia arxiv

Prompt injection attacks pose a serious security risk to large language models in real-world applications. However, existing prompt injection benchmarks primarily focus on short-context inputs, leaving the attacks and de…

A Critical Evaluation of Defenses against Prompt Injection Attacks

2025-05-23 · Yuqi Jia, Zedian Shao, Yupei Liu, Jinyuan Jia 외

Large Language Models (LLMs) are vulnerable to prompt injection attacks, and several defenses have recently been proposed, often claiming to mitigate these attacks successfully. However, we argue that existing studies la…

Formalizing and Benchmarking Prompt Injection Attacks and Defenses

2023-10-19 · Yupei Liu, Yuqi Jia, Runpeng Geng, Jinyuan Jia 외

A prompt injection attack aims to inject malicious instruction/data into the input of an LLM-Integrated Application such that it produces results as an attacker desires. Existing works are limited to case studies. As a r…

Benchmarking

Adaptive Attacks Break Defenses Against Indirect Prompt Injection Attacks on LLM Agents

2025-02-27 · Qiusi Zhan, Richard Fang, Henil Shalin Panchal, Daniel Kang

Large Language Model (LLM) agents exhibit remarkable performance across diverse applications by using external tools to interact with environments. However, integrating external tools introduces security risks, such as i…

Language ModelingLanguage ModellingLarge Language Model

PIArena: A Platform for Prompt Injection Evaluation

2026-04-09 · Runpeng Geng, Chenlong Yin, Yanting Wang, Ying Chen 외 arxiv

Prompt injection attacks pose serious security risks across a wide range of real-world applications. While receiving increasing attention, the community faces a critical gap: the lack of a unified platform for prompt inj…