paper-with-me

홈 › Papers

Reasoning Provenance for Autonomous AI Agents: Structured Behavioral Analytics Beyond State Checkpoints and Execution Traces

2026-03-23 · Neelmani Vispute, Aditya Kadam arxiv

As AI agents transition from human-supervised copilots to autonomous platform infrastructure, the ability to analyze their reasoning behavior across populations of investigations becomes a pressing infrastructure requirement. Existing operational tooling addresses adjacent needs effectively: state checkpoint systems enable fault tolerance; observability platforms provide execution traces for debugging; telemetry standards ensure interoperability. What current systems do not natively provide as a first-class, schema-level primitive is structured reasoning provenance -- normalized, queryable records of why the agent chose each action, what it concluded from each observation, how each conclusion shaped its strategy, and which evidence supports its final verdict. This paper introduces the Agent Execution Record (AER), a structured reasoning provenance primitive that captures intent, observation, and inference as first-class queryable fields on every step, alongside versioned plans with revision rationale, evidence chains, structured verdicts with confidence scores, and delegation authority chains. We formalize the distinction between computational state persistence and reasoning provenance, argue that the latter cannot in general be faithfully reconstructed from the former, and show how AERs enable population-level behavioral analytics: reasoning pattern mining, confidence calibration, cross-agent comparison, and counterfactual regression testing via mock replay. We present a domain-agnostic model with extensible domain profiles, a reference implementation and SDK, and outline an evaluation methodology informed by preliminary deployment on a production platformized root cause analysis agent.

📄 PDF Abstract BibTeX arXiv:2603.21692

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Structured Episodic Event Memory

2026-01-10 · Zhengxuan Lu, Dongfang Li, Yukun Shi, Beilun Wang 외 arxiv

Current approaches to memory in Large Language Models (LLMs) predominantly rely on static Retrieval-Augmented Generation (RAG), which often results in scattered retrieval and fails to capture the structural dependencies …

Autonomous Agents Coordinating Distributed Discovery Through Emergent Artifact Exchange

2026-03-15 · Fiona Y. Wang, Lee Marom, Subhadeep Pal, Rachel K. Luu 외 arxiv

We present ScienceClaw + Infinite, a framework for autonomous scientific investigation in which independent agents conduct research without central coordination, and any contributor can deploy new agents into a shared ec…

LEDGERMIND: Provenance-Constrained Multimodal Agentic Reasoning with a Structured Evidence Ledger

2026-07-30 · Enjun Du, Hange Zhou, Chenxu Du, Siyi Liu 외 arxiv

Multimodal agents for visual question answering increasingly operate as multi-step trajectories that interleave perception, retrieval, and reasoning, yet evaluation still largely reduces to final-answer accuracy. This ag…

Visual Question AnsweringMultimodal Reasoning

Agent-ToM: Learning to Monitor Autonomous LLM Agents via Theory-of-Mind Reasoning

2026-05-22 · Nesreen K. Ahmed, Nima Nafisi arxiv

Monitoring autonomous large language model (LLM) agents for covert malicious behavior is challenging due to delayed, context-dependent, and long-horizon attack patterns. Agents may pursue hidden objectives while maintain…

TRACER: Verifiable Generative Provenance for Multimodal Tool-Using Agents

2026-05-11 · Bihui Yu, Caijun Jia, Jing Chi, Xiaohan Liu 외 arxiv

Multimodal large language models increasingly solve vision-centric tasks by calling external tools for visual inspection, OCR, retrieval, calculation, and multi-step reasoning. Current tool-using agents usually expose th…

Reinforcement Learning