paper-with-me

홈 › Papers

Hardware Keystores for AI Agent Signing Workflows: A Zero-Trust MCP Enforcement Architecture

2026-08-06 · Leo Sambrook, Sampo Sovio arxiv

AI agents performing cryptographic operations (signing Git commits, authenticating API calls, issuing certificates) currently store private keys in software-accessible locations: plaintext files, environment variables, or container memory. Any process with sufficient read privileges can extract the raw key material. A recent production incident demonstrated the practical severity: private keys were exfiltrated from a widely deployed framework via email injection in under five minutes. We aim to enforce both key confidentiality and content-aware authorisation for key use. To that end, we replace software-resident keys with hardware-confined keys accessible through a vendor-neutral PKCS#11 interface. A hardware keystore (HSM, TPM, smart card) executes cryptographic operations on-device; the host receives only the result via opaque handles. Hardware confinement is the primary contribution; it is enabled by a surrounding five-layer Zero-Trust enforcement stack comprising session identity (SAGA), scope bounds (Smax), semantic validation (RAV), taint tracking, and the hardware execution boundary. We evaluate against 12 injection scenarios derived from AgentDojo's ImportantInstructionsAttack template (Debenedetti et al., arXiv:2406.13352). We run four LLM models; three follow injections in baseline mode (gpt-oss-120b, Qwen2.5-72B, DeepSeek-V4-Flash, n=192 combined). Baseline Attack Success Rate (ASR): 19.3% [14.3%, 25.4%]; protected ASR: 0% (Wilson 95% CI upper bound 2.0%). Zero false positives across four benign task scenarios.

📄 PDF Abstract BibTeX arXiv:2608.06130

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

ComfyBench: Benchmarking LLM-based Agents in ComfyUI for Autonomously Designing Collaborative AI Systems

2024-09-02 · CVPR 2025 1 · Xiangyuan Xue, Zeyu Lu, Di Huang, Zidong Wang 외

Much previous AI research has focused on developing monolithic models to maximize their intelligence, with the primary goal of enhancing performance on specific tasks. In contrast, this work attempts to study using LLM-b…

BenchmarkingInstruction Following

AgentCo-op: Retrieval-Based Synthesis of Interoperable Multi-Agent Workflows

2026-05-19 · Shuaike Shen, Wenduo Cheng, Shike Wang, Mingqian Ma 외 arxiv

Designing multi-agent workflows is especially difficult in open-ended scientific settings where tasks lack curated training sets, reliable scalar evaluation metrics, and standardized interfaces between existing tools and…

Murakkab: Resource-Efficient Agentic Workflow Orchestration in Cloud Platforms

2025-08-22 · Gohar Irfan Chaudhry, Esha Choukse, Haoran Qiu, Íñigo Goiri 외 arxiv

Agentic workflows commonly coordinate multiple models and tools with complex control logic. They are quickly becoming the dominant paradigm for AI applications. However, serving them remains inefficient with today's fram…

FlowSteer: Towards Agents Designing Agentic Workflows via Reinforced Progressive Canvas Editing

2026-02-02 · Mingda Zhang, Wenjin Liu, Tiesunlong Shen, Qika Lin 외 arxiv

In recent years, agentic workflows have been widely applied to solve complex human tasks. However, existing workflow construction still faces key challenges, including human-dependent workflow construction, the lack of g…

Reinforcement Learning

Towards Multifaceted Human-Centered AI

2023-01-09 · Sajjadur Rahman, Hannah Kim, Dan Zhang, Estevam Hruschka 외

Human-centered AI workflows involve stakeholders with multiple roles interacting with each other and automated agents to accomplish diverse tasks. In this paper, we call for a holistic view when designing support mechani…