paper-with-me

홈 › Papers

The Steganographic Potentials of Language Models

2025-05-06 · Artem Karpov, Tinuade Adeleke, Seong Hah Cho, Natalia Perez-Campanero

The potential for large language models (LLMs) to hide messages within plain text (steganography) poses a challenge to detection and thwarting of unaligned AI agents, and undermines faithfulness of LLMs reasoning. We explore the steganographic capabilities of LLMs fine-tuned via reinforcement learning (RL) to: (1) develop covert encoding schemes, (2) engage in steganography when prompted, and (3) utilize steganography in realistic scenarios where hidden reasoning is likely, but not prompted. In these scenarios, we detect the intention of LLMs to hide their reasoning as well as their steganography performance. Our findings in the fine-tuning experiments as well as in behavioral non fine-tuning evaluations reveal that while current models exhibit rudimentary steganographic abilities in terms of security and capacity, explicit algorithmic guidance markedly enhances their capacity for information concealment.

📄 PDF Abstract BibTeX arXiv:2505.03439

Code (0)

등록된 구현이 없습니다.

Tasks

Reinforcement Learning (RL)

Methods 이 논문이 사용한 방법론

NON 설명 없음

Similar Papers 제목 키워드 기반

ADLM -- stega: A Universal Adaptive Token Selection Algorithm for Improving Steganographic Text Quality via Information Entropy

2024-10-28 · Zezheng Qin, Congcong Sun, Taiyi He, Yuke He 외

In the context of widespread global information sharing, information security and privacy protection have become focal points. Steganographic systems enhance information security by embedding confidential information int…

DiversityText Generation

A Decision-Theoretic Formalisation of Steganography With Applications to LLM Monitoring

2026-02-26 · Usman Anwar, Julianna Piskorz, David D. Baek, David Africa 외 arxiv

Large language models are beginning to show steganographic capabilities. Such capabilities could allow misaligned models to evade oversight mechanisms. Yet principled methods to detect and quantify such behaviours are la…

NEST: Nascent Encoded Steganographic Thoughts

2026-02-15 · Artem Karpov arxiv

Monitoring chain-of-thought (CoT) reasoning is a foundational safety technique for large language model agents; however, this oversight is compromised if models learn to conceal their reasoning. We explore steganographic…

Purified and Unified Steganographic Network

2024-02-27 · CVPR 2024 1 · Guobiao Li, Sheng Li, Zicong Luo, Zhenxing Qian 외

Steganography is the art of hiding secret data into the cover media for covert communication. In recent years, more and more deep neural network (DNN)-based steganographic schemes are proposed to train steganographic net…

DenoisingImage Denoising

Steganography of Steganographic Networks

2023-02-28 · Guobiao Li, Sheng Li, Meiling Li, Xinpeng Zhang 외

Steganography is a technique for covert communication between two parties. With the rapid development of deep neural networks (DNN), more and more steganographic networks are proposed recently, which are shown to be prom…