paper-with-me

홈 › Papers

Now You (Still) See Me: Detecting Evasive Steganographic Payloads in LLMs

2026-06-08 · Charles Westphal, Timothy Douglas, Keivan Navaie, Tiago Pimentel, Fernando E. Rosas arxiv

Large language models can be fine-tuned to encode prompt-borne secrets into fluent, seemingly benign outputs. This creates a steganographic exfiltration risk that is difficult to detect with output-level steganalysis. Recent work proposes mechanistic detection using linear probes that recover the secret from internal activations. We show that this defense can be systematically evaded, but that detectability can be recovered through a targeted data-level intervention. First, we extend the detection setup to include a non-linear MLP probe. We then adversarially fine-tune steganographic trojans across five base models: Qwen3-8B, Llama-3.1-8B, Ministral-8B, Qwen3-14B, and Phi-4-14B. The resulting models retain $58$--$79\%$ exact-match secret recovery while evading both ridge and held-out MLP probes, with $1$--$8\%$ average capability degradation across six benchmarks. We then give an information-theoretic characterization of this evasion. Successful evasion preserves recoverability while reducing low-order extractability of the secret from the content-aligned representation, forcing the payload into synergistic interaction with residual degrees of freedom. This motivates a recontextualization dataset that restricts these residual degrees of freedom. On this distribution, both ridge and MLP detectability are restored across all five evasive trojans. Overall, our findings show that activation-based steganography detection is vulnerable to adaptive evasion, but also that theory-guided evaluation distributions can expose otherwise hidden payloads.

📄 PDF Abstract BibTeX arXiv:2606.09411

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Steganographic Generative Adversarial Networks

2017-03-16 · Denis Volkhonskiy, Ivan Nazarov, Evgeny Burnaev

Steganography is collection of methods to hide secret information ("payload") within non-secret information "container"). Its counterpart, Steganalysis, is the practice of determining if a message contains a hidden paylo…

Steganalysis

Destruction of Image Steganography using Generative Adversarial Networks

2019-12-20 · Isaac Corley, Jonathan Lwowski, Justin Hoffman

Digital image steganalysis, or the detection of image steganography, has been studied in depth for years and is driven by Advanced Persistent Threat (APT) groups', such as APT37 Reaper, utilization of steganographic tech…

BlockingGenerative Adversarial NetworkImage SteganographySteganalysis+1

Generating Steganographic Text with LSTMs

2017-05-30 · ACL 2017 7 · Tina Fang, Martin Jaggi, Katerina Argyraki

Motivated by concerns for user privacy, we design a steganographic system ("stegosystem") that enables two users to exchange encrypted messages without an adversary detecting that such an exchange is taking place. We pro…

Kolmogorov Complexity Bounds for LLM Steganography and a Perplexity-Based Detection Proxy

2026-03-23 · Andrii Shportko arxiv

Large language models can rewrite text to embed hidden payloads while preserving surface-level meaning, a capability that opens covert channels between cooperating AI systems and poses challenges for alignment monitoring…

A Decision-Theoretic Formalisation of Steganography With Applications to LLM Monitoring

2026-02-26 · Usman Anwar, Julianna Piskorz, David D. Baek, David Africa 외 arxiv

Large language models are beginning to show steganographic capabilities. Such capabilities could allow misaligned models to evade oversight mechanisms. Yet principled methods to detect and quantify such behaviours are la…