paper-with-me

홈 › Papers

A Decision-Theoretic Formalisation of Steganography With Applications to LLM Monitoring

2026-02-26 · Usman Anwar, Julianna Piskorz, David D. Baek, David Africa, Jim Weatherall, Max Tegmark, Christian Schroeder de Witt, Mihaela van der Schaar, David Krueger arxiv

Large language models are beginning to show steganographic capabilities. Such capabilities could allow misaligned models to evade oversight mechanisms. Yet principled methods to detect and quantify such behaviours are lacking. Classical definitions of steganography, and detection methods based on them, require a known reference distribution of non-steganographic signals. For the case of steganographic reasoning in LLMs, knowing such a reference distribution is not feasible; this renders these approaches inapplicable. We propose an alternative, \textbf{decision-theoretic view of steganography}. Our central insight is that steganography creates an asymmetry in usable information between agents who can and cannot decode the hidden content (present within a steganographic signal), and this otherwise latent asymmetry can be inferred from the agents' observable actions. To formalise this perspective, we introduce generalised $\mathcal{V}$-information: a utilitarian framework for measuring the amount of usable information within some input. We use this to define the \textbf{steganographic gap} -- a measure that quantifies steganography by comparing the downstream utility of the steganographic signal to agents that can and cannot decode the hidden content. We empirically validate our formalism, and show that it can be used to detect, quantify, and mitigate steganographic reasoning in LLMs.

📄 PDF Abstract BibTeX arXiv:2602.23163

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Perfectly Secure Steganography Using Minimum Entropy Coupling

2022-10-24 · Christian Schroeder de Witt, Samuel Sokota, J. Zico Kolter, Jakob Foerster 외

Steganography is the practice of encoding secret information into innocuous content in such a manner that an adversarial third party would not realize that there is hidden meaning. While this problem has classically been…

Early Signs of Steganographic Capabilities in Frontier LLMs

2025-07-03 · Artur Zolkowski, Kei Nishimura-Gasparian, Robert McCarthy, Roland S. Zimmermann 외

Monitoring Large Language Model (LLM) outputs is crucial for mitigating risks from misuse and misalignment. However, LLMs could evade monitoring through steganography: Encoding hidden information within seemingly benign …

Large Language Model

A Survey On Semantic Steganography Systems

2022-02-03 · João Figueira

Steganography is the practice of concealing a message within some other carrier or cover message. It is used to allow the sending of hidden information through communication channels where third parties would only be awa…

Survey

Robust Invertible Image Steganography

2022-01-01 · CVPR 2022 1 · Youmin Xu, Chong Mou, Yujie Hu, Jingfen Xie 외

Image steganography aims to hide secret images into a container image, where the secret is hidden from human vision and can be restored when necessary. Previous image steganography methods are limited in hiding capac…

Image Steganography

Secret Collusion among Generative AI Agents: Multi-Agent Deception via Steganography

2024-02-12 · Sumeet Ramesh Motwani, Mikhail Baranchuk, Martin Strohmeier, Vijay Bolina 외

Recent capability increases in large language models (LLMs) open up applications in which groups of communicating generative AI agents solve joint tasks. This poses privacy and security challenges concerning the unauthor…