paper-with-me

홈 › Papers

Safeguarding LLM Agents from Misalignment through Provenance Analysis

2026-05-01 · Yining She, Yiliang Liang, Eunsuk Kang arxiv

As LLM agents gain increasing access to powerful tools, ensuring that their actions are aligned with the user's intent becomes critical. When an agent's proposed tool invocation deviates from the user's intent -- a phenomenon called misalignment -- it may lead to harmful consequences that are difficult to undo. Existing runtime guardrails rely on an LLM-as-a-judge paradigm that lacks a systematic framework for reasoning about alignment, often producing judgments that are inconsistent or difficult to audit. Motivated by provenance analysis, we propose a provenance-based conceptual framework that formalizes misalignment detection as determining whether a proposed tool call is supported by traceable evidence in the agent's context. Building on this framework, we propose ProvenanceGuard, a multi-stage pipeline that analyzes the agent's action for three types of misalignment before the selected tool is executed and only allows the action to take place when it is considered aligned with the user's input query. We evaluated our proposed approach on two different benchmarks, Agent-SafetyBench and WorkBench, across 10 backbone LLMs. Compared to the LLM-as-a-judge baseline, ProvenanceGuard reduces error rate on misaligned traces from 42.9% to 1.8% on Agent-SafetyBench and from 32.1% to 17.3% on WorkBench, while reducing intervention burden on task-successful traces from 30.5% to 12.8% and introducing no statistically significant increase in unnecessary interventions on aligned traces. These results demonstrate that structured, provenance-based reasoning provides an effective and practical foundation for safeguarding LLM agents from misalignment.

📄 PDF Abstract BibTeX arXiv:2607.01236

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Provenance-Based Interpretation of Multi-Agent Information Analysis

2020-11-08 · Scott Friedman, Jeff Rye, David LaVergne, Dan Thomsen 외

Analytic software tools and workflows are increasing in capability, complexity, number, and scale, and the integrity of our workflows is as important as ever. Specifically, we must be able to inspect the process of analy…

Diversity

Causality Laundering: Denial-Feedback Leakage in Tool-Calling LLM Agents

2026-04-05 · Mohammad Hossein Chinaei arxiv

Tool-calling LLM agents can read private data, invoke external services, and trigger real-world actions, creating a security problem at the point of tool execution. We identify a denial-feedback leakage pattern, which we…

LLM Agents for Interactive Workflow Provenance: Reference Architecture and Evaluation Methodology

2025-09-17 · Renan Souza, Timothy Poteet, Brian Etz, Daniel Rosendo 외 arxiv

Modern scientific discovery increasingly relies on workflows that process data across the Edge, Cloud, and High Performance Computing (HPC) continuum. Comprehensive and in-depth analyses of these data are critical for hy…

Anomaly Detection

How Coding Agents Fail Their Users: A Large-Scale Analysis of Developer-Agent Misalignment in 20,574 Real-World Sessions

2026-05-28 · Ningzhi Tang, Chaoran Chen, Gelei Xu, Yiyu Shi 외 arxiv

AI coding agents increasingly act directly within software environments, yet existing analyses of their failures rely on benchmark trajectories that miss how developers actually experience misalignment. We present an obs…

TRACER: Verifiable Generative Provenance for Multimodal Tool-Using Agents

2026-05-11 · Bihui Yu, Caijun Jia, Jing Chi, Xiaohan Liu 외 arxiv

Multimodal large language models increasingly solve vision-centric tasks by calling external tools for visual inspection, OCR, retrieval, calculation, and multi-step reasoning. Current tool-using agents usually expose th…

Reinforcement Learning