paper-with-me

홈 › Papers

Prompt Injection Mitigation with Agentic AI, Nested Learning, and AI Sustainability via Semantic Caching

2026-01-19 · Diego Gosmar, Deborah A. Dahl arxiv

Prompt injection remains a central obstacle to the safe deployment of large language models, particularly in multi-agent settings where intermediate outputs can propagate or amplify malicious instructions. Building on earlier work that introduced a four-metric Total Injection Vulnerability Score (TIVS), this paper extends the evaluation framework with semantic similarity-based caching and a fifth metric (Observability Score Ratio) to yield TIVS-O, investigating how defence effectiveness interacts with transparency in a HOPE-inspired Nested Learning architecture. The proposed system combines an agentic pipeline with Continuum Memory Systems that implement semantic similarity-based caching across 301 synthetically generated injection-focused prompts drawn from ten attack families, while a fourth agent performs comprehensive security analysis using five key performance indicators. In addition to traditional injection metrics, OSR quantifies the richness and clarity of security-relevant reasoning exposed by each agent, enabling an explicit analysis of trade-offs between strict mitigation and auditability. Experiments show that the system achieves secure responses with zero high-risk breaches, while semantic caching delivers substantial computational savings, achieving a 41.6% reduction in LLM calls and corresponding decreases in latency, energy consumption, and carbon emissions. Five TIVS-O configurations reveal optimal trade-offs between mitigation strictness and forensic transparency. These results indicate that observability-aware evaluation can reveal non-monotonic effects within multi-agent pipelines and that memory-augmented agents can jointly maximize security robustness, real-time performance, operational cost savings, and environmental sustainability without modifying underlying model weights, providing a production-ready pathway for secure and green LLM deployments.

📄 PDF Abstract BibTeX arXiv:2601.13186

Code (0)

등록된 구현이 없습니다.

Tasks

Semantic Similarity

Similar Papers 제목 키워드 기반

Hallucination Mitigation with Agentic AI, Nested Learning, and AI Sustainability via Semantic Caching

2026-05-27 · Diego Gosmar, Deborah A. Dahl arxiv

Hallucination remains a major reliability barrier for production LLM systems, particularly in multi-agent pipelines where unsupported claims can propagate unchecked across stages. This paper adapts a HOPE-inspired Nested…

Semantic Similarity

QSAF: A Novel Mitigation Framework for Cognitive Degradation in Agentic AI

2025-07-21 · Hammad Atta, Muhammad Zeeshan Baig, Yasir Mehmood, Nadeem Shahzad 외 arxiv

We introduce Cognitive Degradation as a novel vulnerability class in agentic AI systems. Unlike traditional adversarial external threats such as prompt injection, these failures originate internally, arising from memory …

What If Prompt Injection Never Left? Exploring Cross-Session Stored Prompt Injection in Agentic Systems

2026-06-03 · Yuanbo Xie, Tianyun Liu, Yingjie Zhang, Suchen Liu 외 arxiv

Modern agentic systems transform LLMs from session-bounded assistants into stateful systems that persist and evolve shared world state across sessions through memories, filesystems, tools, and other long-lived contextual…

Meta SecAlign: A Secure Foundation LLM Against Prompt Injection Attacks

2025-07-03 · Sizhe Chen, Arman Zharmagambetov, David Wagner, Chuan Guo

Prompt injection attacks pose a significant security threat to LLM-integrated applications. Model-level defenses have shown strong effectiveness, but are currently deployed into commercial-grade models in a closed-source…

Instruction Following

Prompt Injection 2.0: Hybrid AI Threats

2025-07-17 · Jeremy McHugh, Kristina Šekrst, Jon Cefalu

Prompt injection attacks, where malicious input is designed to manipulate AI systems into ignoring their original instructions and following unauthorized commands instead, were first discovered by Preamble, Inc. in May 2…