paper-with-me

Papers

Hiding in Plain Floats: Steganographic Carriers for Indirect Prompt and Content Injection

2026-06-07 · Mudit Sinha, Sanika Chavan arxiv

Text-centered prompt-injection defenses assume that the malicious signal is visible in one of the inspected text views. We study a reproducible LLM01-style indirect prompt/content-injection failure mode where that assumption breaks: a payload caught in plain English slips past the same detector when it is transported as structured float parameters and reconstructed only as fragmented telemetry. Across 14,400 attacked real-model trials on three commercial LLM APIs from different providers, the IFS-derived float-array carrier preserves 94.3% leakage ASR under the strongest dual-layer text-classifier defense evaluated in the main matrix: a Prompt Guard 2 + TF-IDF ensemble; the same carrier-level pattern also replicates with a fine-tuned roberta-base detector. We emphasize leakage ASR because downstream systems may act on quoted or reproduced markers even when the model refuses, but Strong ASR is the stricter metric for structurally compliant attack success. A 2 x 2 ablation shows that data-layer storage and reconstruction-layer fragmentation defeat different text views and that both are needed to evade both. A simple xxd detector and semantic validation block the current T3 instance, so the contribution is not an undetectable exploit but a measured failure boundary for text-only inspection in structured-input pipelines that expose reconstructed auxiliary channels to an LLM.

📄 PDF Abstract BibTeX arXiv:2606.08403

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Hidden in Plain Text: Emergence & Mitigation of Steganographic Collusion in LLMs

2024-10-02 · Yohan Mathew, Ollie Matthews, Robert McCarthy, Joan Velja 외

The rapid proliferation of frontier model agents promises significant societal advances but also raises concerns about systemic risks arising from unsafe interactions. Collusion to the disadvantage of others has been ide…

In-Context Reinforcement Learningreinforcement-learningReinforcement Learning

SUDS: Sanitizing Universal and Dependent Steganography

2023-09-23 · Preston K. Robinette, Hanchen D. Wang, Nishan Shehadeh, Daniel Moyer 외

Steganography, or hiding messages in plain sight, is a form of information hiding that is most commonly used for covert communication. As modern steganographic mediums include images, text, audio, and video, this communi…

Steganalysis

Hiding Information in Big Data based on Deep Learning

2019-12-31 · Dingju Zhu

The current approach of information hiding based on deep learning model can not directly use the original data as carriers, which means the approach can not make use of the existing data in big data to hiding information…

Deep Learning

$\mathbf{S^2LM}$: Towards Semantic Steganography via Large Language Models

2025-11-07 · Huanqi Wu, Huangbiao Xu, Runfeng Xie, Jiaxin Cai 외 arxiv

Despite remarkable progress in steganography, embedding semantically rich, sentence-level information into carriers remains a challenging problem. In this work, we present a novel concept of Semantic Steganography, which…

Hiding Images in Plain Sight: Deep Steganography

2017-12-01 · NeurIPS 2017 12 · Shumeet Baluja

Steganography is the practice of concealing a secret message within another, ordinary, message. Commonly, steganography is used to unobtrusively hide a small message within the noisy regions of a larger image. In this …