paper-with-me

홈 › Papers

PISmith: Reinforcement Learning-based Red Teaming for Prompt Injection Defenses

2026-03-13 · Chenlong Yin, Runpeng Geng, Yanting Wang, Jinyuan Jia arxiv

Prompt injection poses serious security risks to real-world LLM applications, particularly autonomous agents. Although many defenses have been proposed, their robustness against adaptive attacks remains insufficiently evaluated, potentially creating a false sense of security. In this work, we propose PISmith, a reinforcement learning (RL)-based red-teaming framework that systematically assesses existing prompt-injection defenses by training an attack LLM to optimize injected prompts in a practical black-box setting, where the attacker can only query the defended LLM and observe its outputs. We find that directly applying standard GRPO to attack strong defenses leads to sub-optimal performance due to extreme reward sparsity -- most generated injected prompts are blocked by the defense, causing the policy's entropy to collapse before discovering effective attack strategies, while the rare successes cannot be learned effectively. In response, we introduce adaptive entropy regularization and dynamic advantage weighting to sustain exploration and amplify learning from scarce successes. Extensive evaluation on 13 benchmarks demonstrates that state-of-the-art prompt injection defenses remain vulnerable to adaptive attacks. We also compare PISmith with 7 baselines across static, search-based, and RL-based attack categories, showing that PISmith consistently achieves the highest attack success rates. Furthermore, PISmith achieves strong performance in agentic settings on InjecAgent and AgentDojo against both open-source and closed-source LLMs (e.g., GPT-4o-mini and GPT-5-nano). Our code is available at https://github.com/albert-y1n/PISmith.

📄 PDF Abstract BibTeX arXiv:2603.13026

Code (0)

등록된 구현이 없습니다.

Tasks

Reinforcement LearningRed Teaming

Similar Papers 제목 키워드 기반

RL Is a Hammer and LLMs Are Nails: A Simple Reinforcement Learning Recipe for Strong Prompt Injection

2025-10-06 · Yuxin Wen, Arman Zharmagambetov, Ivan Evtimov, Narine Kokhlikyan 외 arxiv

Prompt injection poses a serious threat to the reliability and safety of LLM agents. Recent defenses against prompt injection, such as Instruction Hierarchy and SecAlign, have shown notable robustness against static atta…

Reinforcement Learning

Agent Against Agent: An Agentic System for Automatic Prompt Injection Red Teaming

2026-08-05 · Yanting Wang, Chenlong Yin, Runpeng Geng, Jinyuan Jia hf

Prompt injection poses significant security risks to LLM agents. Efficient and effective red-teaming is therefore critical, both for evaluating these risks and for collecting training data to improve defenses. Existing s…

Reinforcement LearningRed Teaming

PI-Hunter: Automated Red-Teaming for Exposing and Localizing Prompt Injections

2026-06-10 · Pengfei He, Lesly Miculicich, Vishesh Sharma, Ash Fox 외 arxiv

Large Language Models (LLMs) are rapidly evolving into agentic systems that interact with external tools and environments, introducing new security risks such as indirect prompt injection attacks through untrusted extern…

Recent advancements in LLM Red-Teaming: Techniques, Defenses, and Ethical Considerations

2024-10-09 · Tarun Raheja, Nilay Pochhi, F. D. C. M. Curie

Large Language Models (LLMs) have demonstrated remarkable capabilities in natural language processing tasks, but their vulnerability to jailbreak attacks poses significant security risks. This survey paper presents a com…

Language ModelingLanguage ModellingLarge Language ModelPrompt Engineering+1

Defending against Adaptive Prompt Injection Attacks via Reasoning-enabled Task Alignment

2026-06-13 · Lipeng He, Yihan Wang, Jiawen Zhang, N. Asokan arxiv

Indirect prompt injection attacks hijack LLM-based agents by embedding malicious instructions in third-party data that the agent retrieves during task execution. Existing defenses report near-zero attack success rate on …

Reinforcement Learning