paper-with-me

홈 › Papers

SafeHarness: Lifecycle-Integrated Security Architecture for LLM-based Agent Deployment

2026-04-15 · Xixun Lin, Yang Liu, Yancheng Chen, Yongxuan Wu, Yucheng Ning, Yilong Liu, Nan Sun, Shun Zhang, Bin Chong, Chuan Zhou, Yanan Cao arxiv

The performance of large language model (LLM) agents depends critically on the execution harness, the system layer that orchestrates tool use, context management, and state persistence. Yet this same architectural centrality makes the harness a high-value attack surface: a single compromise at the harness level can cascade through the entire execution pipeline. We observe that existing security approaches suffer from structural mismatch, leaving them blind to harness-internal state and unable to coordinate across the different phases of agent operation. In this paper, we introduce \safeharness{}, a security architecture in which four proposed defense layers are woven directly into the agent lifecycle to address above significant limitations: adversarial context filtering at input processing, tiered causal verification at decision making, privilege-separated tool control at action execution, and safe rollback with adaptive degradation at state update. The proposed cross-layer mechanisms tie these layers together, escalating verification rigor, triggering rollbacks, and tightening tool privileges whenever sustained anomalies are detected. We evaluate \safeharness{} on benchmark datasets across diverse harness configurations, comparing against four security baselines under five attack scenarios spanning six threat categories. Compared to the unprotected baseline, \safeharness{} achieves an average reduction of approximately 38\% in UBR and 42\% in ASR, substantially lowering both the unsafe behavior rate and the attack success rate while preserving core task utility.

📄 PDF Abstract BibTeX arXiv:2604.13630

Code (0)

등록된 구현이 없습니다.

Tasks

Decision Making

Similar Papers 제목 키워드 기반

LLM agents security duality: a comprehensive survey of self-security and empowered cybersecurity

2026-06-26 · Yiwei Xu, Yong Zhuang, Xuanming Liu, Tian Zhang 외 arxiv

Large language model (LLM) agents are rapidly being integrated into real-world systems. Their autonomy and tool-use capabilities generate substantial value while simultaneously expanding the security attack surface. This…

AgentKernel: The Trust-Native Agentic Operating System

2026-08-29 · Zhenhua Zou, Sheng Guo, Qiuyang Zhan, Lepeng Zhao 외 hf

Modern AI agents routinely cross trust boundaries: they ingest untrusted content, combine it with privileged instructions, persist intermediate beliefs in long-term memory, and invoke privileged tools. This creates an at…

Taming OpenClaw: Security Analysis and Mitigation of Autonomous LLM Agent Threats

2026-03-12 · Xinhao Deng, Yixiang Zhang, Jiaqing Wu, Jiaqi Bai 외 arxiv

Autonomous Large Language Model (LLM) agents, exemplified by OpenClaw, demonstrate remarkable capabilities in executing complex, long-horizon tasks. However, their tightly coupled instant-messaging interaction paradigm a…

A Survey on Agentic Security: Applications, Threats and Defenses

2025-10-07 · Asif Shahriar, Md Nafiu Rahman, Sadif Ahmed, Farig Sadeque 외 arxiv

LLM-based agents are now used throughout cybersecurity. While these agents facilitate powerful and autonomous security applications, their autonomy opens up new attack surfaces, and the security community is actively bui…

AgentWard: A Lifecycle Security Architecture for Autonomous AI Agents

2026-04-27 · Yixiang Zhang, Xinhao Deng, Jiaqing Wu, Yue Xiao 외 arxiv

Autonomous AI agents extend large language models into full runtime systems that load skills, ingest external content, maintain memory, plan multi-step actions, and invoke privileged tools. In such systems, security fail…