paper-with-me

Papers

Document-Authored Control-Signal Impersonation: A Low-Cost Indirect Prompt Attack on RAG Safety Boundaries

2026-06-08 · Jianguo Zhu arxiv

Retrieval-augmented generation (RAG) systems often serialize user queries, retrieved documents, metadata, system labels, and task instructions into one natural-language prompt. We study a source-authority boundary failure in this design: attacker-authored retrieved text can impersonate metadata, provenance, authority, or disclosure-policy signals that appear control-relevant to the model. We call this pattern Document-Authored Control-Signal Impersonation (DACSI). DACSI is a non-imperative, metadata-like payload subclass within indirect prompt injection. Its central lesson is simple: document-authored labels are data, not policy. Command-style injection asks the model to ignore, override, or violate policy; DACSI asks whether untrusted document text can be misattributed as an authorized control signal when RAG prompt rendering collapses trusted and untrusted text into the same natural-language channel. We evaluate DACSI across six model settings, prompt-pressure levels, injection baselines, signal taxonomies, RAG-mediated pipelines, system-control probes, a source-authority attribution probe, and synthetic canary formats. We interpret the evidence by model regime rather than as six equal replications: DeepSeek V4 Pro and Qwen3.5-397B provide the cleanest positive lift, DeepSeek V4 Flash is a high-susceptibility setting, GPT-5.5 and Gemini 3.1 Pro Low are strong-boundary probes with selected residual risks, and GLM-4.7 is a saturated leakage boundary case. Across these regimes, DACSI warrants separate evaluation because it uses a command-free metadata/provenance/policy surface, follows a RAG-specific source-authority path, and responds to source/channel separation. The source-authority probe is behavioral attribution evidence, not proof of an internal mechanism.

📄 PDF Abstract BibTeX arXiv:2606.09005

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Trading Human Curation for Synthetic Augmentation in RLVR

2026-06-02 · Akshansh, Leonardo Rosa Rodrigues, Michael Korostelev, Youssef Hassan 외 arxiv

The supply of high-quality training tasks is a central bottleneck for reinforcement learning from verifiable rewards (RLVR) on agentic language models. Each task requires a sandboxed setup, a prompt, and a hand-authored …

Reinforcement LearningInstruction Following

Intelligent Documentation in Medical Education: Can AI Replace Manual Case Logging?

2026-01-19 · Nafiz Imtiaz Khan, Kylie Cleland, Vladimir Filkov, Roger Eric Goldman arxiv

Procedural case logs are a core requirement in radiology training, yet they are time-consuming to complete and prone to inconsistency when authored manually. This study investigates whether large language models (LLMs) c…

Transferable Adversarial Face Attack with Text Controlled Attribute

2024-12-16 · Wenyun Li, Zheng Zhang, Xiangyuan Lan, Dongmei Jiang

Traditional adversarial attacks typically produce adversarial examples under norm-constrained conditions, whereas unrestricted adversarial examples are free-form with semantically meaningful perturbations. Current unrest…

AttributeFace Recognition

Stylometry Analysis of Multi-authored Documents for Authorship and Author Style Change Detection

2024-01-12 · Muhammad Tayyab Zamir, Muhammad Asif Ayub, Asma Gul, Nasir Ahmad 외

In recent years, the increasing use of Artificial Intelligence based text generation tools has posed new challenges in document provenance, authentication, and authorship detection. However, advancements in stylometry ha…

Change DetectionStyle change detectionText Generation

Masks and Mimicry: Strategic Obfuscation and Impersonation Attacks on Authorship Verification

2025-03-24 · Kenneth Alperin, Rohan Leekha, Adaku Uchendu, Trang Nguyen 외

The increasing use of Artificial Intelligence (AI) technologies, such as Large Language Models (LLMs) has led to nontrivial improvements in various tasks, including accurate authorship identification of documents. Howeve…

Adversarial RobustnessAuthorship Verification