paper-with-me

홈 › Papers

Unvalidated Trust: Cross-Stage Vulnerabilities in Large Language Model Architectures

2025-10-30 · Dominik Schwarz arxiv

As Large Language Models (LLMs) are increasingly integrated into automated, multi-stage pipelines, risk patterns that arise from unvalidated trust between processing stages become a practical concern. This paper presents a mechanism-centered taxonomy of 41 recurring risk patterns in commercial LLMs. The analysis shows that inputs are often interpreted non-neutrally and can trigger implementation-shaped responses or unintended state changes even without explicit commands. We argue that these behaviors constitute architectural failure modes and that string-level filtering alone is insufficient. To mitigate such cross-stage vulnerabilities, we recommend zero-trust architectural principles, including provenance enforcement, context sealing, and plan revalidation, and we introduce "Countermind" as a conceptual blueprint for implementing these defenses.

📄 PDF Abstract BibTeX arXiv:2510.27190

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Code-Augur: Agentic Vulnerability Detection via Specification Inference

2026-06-17 · Zhengxiong Luo, Mehtab Zafar, Dylan Wolff, Abhik Roychoudhury arxiv

The advent of agentic vulnerability detection is already becoming a watershed moment for software security. Audits conducted entirely by autonomous LLM agents are uncovering critical vulnerabilities in fundamental softwa…

Vulnerability Detection

A Lifecycle and Application-Stack Survey of Large Language Model Vulnerabilities: Attacks, Risks, Defenses, and Open Problems

2026-06-30 · Seyed Bagher Hashemi Natanzi, Bo Tang arxiv

Large language models are no longer only text generators. They are increasingly embedded in retrieval pipelines, enterprise assistants, coding environments, robotic systems, security-operation workflows, and autonomous a…

Red Teaming

Sequential Data Poisoning in LLM Post-Training

2026-06-03 · Jack Sanderson, Yihan Wang, Xiaoqian Lu, Gautam Kamath 외 arxiv

LLM post-training proceeds through multiple stages, e.g., supervised fine-tuning (SFT) followed by reinforcement learning from human feedback (RLHF) or direct preference optimization (DPO), where each stage draws data fr…

Reinforcement Learning

The Computational Drug Repositioning without Negative Sampling

2021-11-29 · Xinxing Yang, Genke Yang, Jian Chu

Computational drug repositioning technology is an effective tool to accelerate drug development. Although this technique has been widely used and successful in recent decades, many existing models still suffer from multi…

Ev-Trust: An Evolutionarily Stable Trust Mechanism for Decentralized LLM-Based Multi-Agent Service Economies

2025-12-18 · Jiye Wang, Shiduo Yang, Ting Qiao, Jiayu Qin 외 arxiv

Decentralized LLM-based multi-agent service economies face three vulnerabilities that undermine traditional trust mechanisms: reduced cost of fraud, difficulty in evaluating service quality, and instability of service co…