paper-with-me

홈 › Papers

Detecting Sleeper Agents in Large Language Models via Semantic Drift Analysis

2025-11-20 · Shahin Zanbaghi, Ryan Rostampour, Farhan Abid, Salim Al Jarmakani arxiv

Large Language Models (LLMs) can be backdoored to exhibit malicious behavior under specific deployment conditions while appearing safe during training a phenomenon known as "sleeper agents." Recent work by Hubinger et al. demonstrated that these backdoors persist through safety training, yet no practical detection methods exist. We present a novel dual-method detection system combining semantic drift analysis with canary baseline comparison to identify backdoored LLMs in real-time. Our approach uses Sentence-BERT embeddings to measure semantic deviation from safe baselines, complemented by injected canary questions that monitor response consistency. Evaluated on the official Cadenza-Labs dolphin-llama3-8B sleeper agent model, our system achieves 92.5% accuracy with 100% precision (zero false positives) and 85% recall. The combined detection method operates in real-time (<1s per query), requires no model modification, and provides the first practical solution to LLM backdoor detection. Our work addresses a critical security gap in AI deployment and demonstrates that embedding-based detection can effectively identify deceptive model behavior without sacrificing deployment efficiency.

📄 PDF Abstract BibTeX arXiv:2511.15992

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

The Trigger in the Haystack: Extracting and Reconstructing LLM Backdoor Triggers

2026-02-03 · Blake Bullwinkel, Giorgio Severi, Keegan Hines, Amanda Minnich 외 arxiv

Detecting whether a model has been poisoned is a longstanding problem in AI security. In this work, we present a practical scanner for identifying sleeper agent-style backdoors in causal language models. Our approach rel…

Plant, Persist, Trigger: Sleeper Attack on Large Language Model Agents

2026-05-27 · Yongxiang Li, Moxin Li, Zhixin Ma, Fengbin Zhu 외 arxiv

Large Language Model (LLM) agents remain vulnerable to safety threats from the external environment, where attackers inject adversarial content into external observations such as tool-returned data, webpages, or MCP cont…

DynaTrust: Defending Multi-Agent Systems Against Sleeper Agents via Dynamic Trust Graphs

2026-03-09 · Yu Li, Qiang Hu, Yao Zhang, Lili Quan 외 arxiv

Large Language Model-based Multi-Agent Systems (MAS) have demonstrated remarkable collaborative reasoning capabilities but introduce new attack surfaces, such as the sleeper agent, which behave benignly during routine op…

Hidden in Memory: Sleeper Memory Poisoning in LLM Agents

2026-05-14 · Sidharth Pulipaka, Stanislau Hlebik, Leonidas Raghav, Sahar Abdelnabi 외 arxiv

Large language models are increasingly augmented with persistent memory, allowing assistants to store user-specific information across sessions for personalization and continuity. This statefulness introduces a new secur…

Sleeper Social Bots: a new generation of AI disinformation bots are already a political threat

2024-08-07 · Jaiv Doshi, Ines Novacic, Curtis Fletcher, Mats Borges 외

This paper presents a study on the growing threat of "sleeper social bots," AI-driven social bots in the political landscape, created to spread disinformation and manipulate public opinion. We based the name sleeper soci…