paper-with-me

홈 › Papers

Latent Adversarial Detection: Adaptive Probing of LLM Activations for Multi-Turn Attack Detection

2026-04-30 · Prashant Kulkarni arxiv

Multi-turn prompt injection follows a known attack path -- trust-building, pivoting, escalation but text-level defenses miss covert attacks where individual turns appear benign. We show this attack path leaves an activation-level signature in the model's residual stream: each phase shift moves the activation, producing a total path length far exceeding benign conversations. We call this adversarial restlessness. Five scalar trajectory features capturing this signal lift conversation-level detection from 76.2% to 93.8% on synthetic held-out data. The signal replicates across four model families (24B-70B); probes are model-specific and do not transfer across architectures. Generalization is source-dependent: leave-one-source-out evaluation shows each of synthetic, LMSYS-Chat-1M, and SafeDialBench captures distinct attack distributions, with detection on real-world LMSYS reaching 47-71% when its distribution is represented in training. Combined three-source training achieves 89.4% detection at 2.4% false positive rate on a held-out mixed set. We further show that three-phase turn-level labels(benign/pivoting/adversarial) unique to our synthetic dataset are essential: binary conversation-level labels produce 50-59% false positives. These results establish adversarial restlessness as a reliable activation-level signal and characterize the data requirements for practical deployment.

📄 PDF Abstract BibTeX arXiv:2604.28129

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Obfuscated Activations Bypass LLM Latent-Space Defenses

2024-12-12 · Luke Bailey, Alex Serrano, Abhay Sheshadri, Mikhail Seleznyov 외

Recent latent-space monitoring techniques have shown promise as defenses against LLM attacks. These defenses act as scanners that seek to detect harmful activations before they lead to undesirable actions. This prompts t…

Probing Latent Subspaces in LLM for AI Security: Identifying and Manipulating Adversarial States

2025-03-12 · Xin Wei Chia, Jonathan Pan

Large Language Models (LLMs) have demonstrated remarkable capabilities across various tasks, yet they remain vulnerable to adversarial manipulations such as jailbreaking via prompt injection attacks. These attacks bypass…

Dimensionality Reduction

Universal Backdoor Attacks Detection via Adaptive Adversarial Probe

2022-09-12 · Yuhang Wang, Huafeng Shi, Rui Min, Ruijia Wu 외

Extensive evidence has demonstrated that deep neural networks (DNNs) are vulnerable to backdoor attacks, which motivates the development of backdoor attacks detection. Most detection methods are designed to verify whethe…

Scheduling

Segment-Level Coherence for Robust Harmful Intent Probing in LLMs

2026-04-16 · Xuanli He, Bilgehan Sel, Faizan Ali, Jenny Bao 외 arxiv

Large Language Models (LLMs) are increasingly exposed to adaptive jailbreaking, particularly in high-stakes Chemical, Biological, Radiological, and Nuclear (CBRN) domains. Although streaming probes enable real-time monit…

Behavior-Aware and Generalizable Defense Against Black-Box Adversarial Attacks for ML-Based IDS

2025-12-15 · Sabrine Ennaji, Elhadj Benkhelifa, Luigi Vincenzo Mancini arxiv

Machine learning based intrusion detection systems are increasingly targeted by black box adversarial attacks, where attackers craft evasive inputs using indirect feedback such as binary outputs or behavioral signals lik…

Change Point DetectionIntrusion DetectionAdversarial Attack