paper-with-me

홈 › Papers

Rewriting the Response Path: Silent Tampering and Provider-Signed Defense in BYOK LLM Agents

2026-05-04 · Mingyu Luo, Zihan Zhang, Zesen Liu, Yuchong Xie, Zhixiang Zhang, Dung Hiu Hilton Yeung, Wai Ip Lai, Ping Chen, Ming Wen, Dongdong She arxiv

LLM agents convert model outputs into consequential actions, including communications, code changes, and financial transactions. Developers often trust evidence such as test results and execution logs. We identify a response path integrity gap in Bring Your Own Key configurations used by roughly 88 percent of mainstream agents. Because traffic passes through a user-authorized relay, the relay can modify plaintext LLM responses after alignment but before execution without breaking encryption. A minimal attack rewrites one execution bearing field and regenerates the remaining response using the user key while preserving the model style. Experiments reveal false green verification, where malicious code modifications pass public tests while silently defeating security checks. On APPS, 99.7 percent of publicly passing solutions retained downgraded behavior without developer-visible warnings. Tests on SWE bench, AgentDojo, and ASB across five frontier models show that single-field rewriting can redirect agents while preserving apparent task completion. We propose sign-c, a server-side scheme that authenticates execution bearing fields and outgoing queries. A local shim verifies them before action, while encryption protects confidentiality. The defense rejected all tampered responses with zero false rejections and only 0.0167 percent latency overhead.

📄 PDF Abstract BibTeX arXiv:2605.02187

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Bridging the Skills Gap: Evaluating an AI-Assisted Provider Platform to Support Care Providers with Empathetic Delivery of Protocolized Therapy

2024-01-08 · William R. Kearns, Jessica Bertram, Myra Divina, Lauren Kemp 외

Despite the high prevalence and burden of mental health conditions, there is a global shortage of mental health providers. Artificial Intelligence (AI) methods have been proposed as a way to address this shortage, by sup…

Sycophancy to Subterfuge: Investigating Reward-Tampering in Large Language Models

2024-06-14 · Carson Denison, Monte MacDiarmid, Fazl Barez, David Duvenaud 외

In reinforcement learning, specification gaming occurs when AI systems learn undesired behaviors that are highly rewarded due to misspecified training goals. Specification gaming can range from simple behaviors like syco…

Language ModellingLarge Language Model

Alignment Tampering: How Reinforcement Learning from Human Feedback Is Exploited to Optimize Misaligned Biases

2026-05-26 · Dongyoon Hahm, Dylan Hadfield-Menell, Kimin Lee arxiv

Reinforcement Learning from Human Feedback (RLHF) is the standard method to align Large Language Models (LLMs) with human preferences. In this work, we introduce alignment tampering, a potential vulnerability where the L…

Reinforcement Learning

Verified Misguidance: Measuring Structural Citation Failures in Search-Augmented LLMs

2026-05-27 · Yongsik Seo, Wooseok Jeong, Eunyoung Kim, Hyeonseo Jang 외 arxiv

Users of search-augmented LLMs rely on citations as evidence that responses are grounded in real sources, and rarely verify the cited pages themselves. Millions of queries per day now pass through these systems, making c…

VERVE: Template-based ReflectiVE Rewriting for MotiVational IntErviewing

2023-11-14 · Do June Min, Verónica Pérez-Rosas, Kenneth Resnicow, Rada Mihalcea

Reflective listening is a fundamental skill that counselors must acquire to achieve proficiency in motivational interviewing (MI). It involves responding in a manner that acknowledges and explores the meaning of what the…