paper-with-me

홈 › Papers

RouteGuard: Internal-Signal Detection of Skill Poisoning in LLM Agents

2026-04-24 · Wenjie Xiao, Xuehai Tang, Biyu Zhou, Songlin Hu, Jizhong Han arxiv

Agent skills introduce a new and more severe form of indirect injection for LLM agents: unlike traditional indirect prompt injection, attackers can hide malicious instructions inside a dense, action-oriented skill that already functions as a legitimate instruction source. We study pre-execution skill-poison detection and show that successful skill poisoning induces a structured internal effect, attention hijacking, in which response-time attention shifts from trusted context to malicious skill spans and drives harmful behavior. Motivated by this mechanism, we propose RouteGuard, a frozen-backbone detector that combines response-conditioned attention and hidden-state alignment through reliability-gated late fusion. Across both real and synthetic open-source skill benchmarks, RouteGuard is consistently the strongest or most robust detector; on the critical Skill-Inject channel slice, it reaches 0.8834 F1 and recovers 90.51% of description attacks missed by lexical screening, showing that defending against skill poisoning requires internal-signal detection rather than text-only filtering

📄 PDF Abstract BibTeX arXiv:2604.22888

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Triaging Threats to Specialized Guardrails

2026-05-29 · Wenjie Jacky Mo, Xiaofei Wen, Rui Cai, Boyu Zhu 외 arxiv

Building robust safety guardrails is essential for deploying Large Language Models across diverse real-world applications. However, this goal remains challenging because safety risks span heterogeneous threat domains, wh…

SKILLC: Learning Autonomous Skill Internalization in LLM Agents via Contrastive Credit Assignment

2026-05-27 · Hongxiang Lin, Zhirui Kuai, Erpeng Xue, Lei Wang arxiv

Structured skill prompts improve exploration in long-horizon agentic reinforcement learning (RL). Skill-augmented RL methods retain external skills at inference, while skill-internalization RL methods withdraw them durin…

Reinforcement Learning

Behavior-Aware and Generalizable Defense Against Black-Box Adversarial Attacks for ML-Based IDS

2025-12-15 · Sabrine Ennaji, Elhadj Benkhelifa, Luigi Vincenzo Mancini arxiv

Machine learning based intrusion detection systems are increasingly targeted by black box adversarial attacks, where attackers craft evasive inputs using indirect feedback such as binary outputs or behavioral signals lik…

Change Point DetectionIntrusion DetectionAdversarial Attack

When Context Bites: Detecting RAG Poisoning via Document-Level Attention Collapse

2026-08-07 · Yingtao Ren, Ziyi Zhao, Yiwei Fu, Xiao Luo 외 hf

Retrieval-augmented generation (RAG) is indispensable for enhancing large language models. However, RAGs are increasingly susceptible to poisoning attacks, in which adversarial documents are injected to manipulate genera…

NeuroGenPoisoning: Neuron-Guided Attacks on Retrieval-Augmented Generation of LLM via Genetic Optimization of External Knowledge

2025-10-24 · Hanyu Zhu, Lance Fiondella, Jiawei Yuan, Kai Zeng 외 arxiv

Retrieval-Augmented Generation (RAG) empowers Large Language Models (LLMs) to dynamically integrate external knowledge during inference, improving their factual accuracy and adaptability. However, adversaries can inject …