paper-with-me

Papers

Trace Integrity for LLM Data Agents: A Vision for Auditable Structured Reasoning in Real-World Systems

2026-08-26 · Srimonti Dutta, Akshata Kishore Moharir arxiv

Answer accuracy is an insufficient reliability signal for LLM data agents. In structured-data tasks, a benchmark-correct answer can be produced by an invalid trace. This paper introduces Trace Integrity, a deployment reliability criterion for evaluating whether the computation recorded behind an answer is explicit, executable, schema-valid, operator-faithful, replayable, answer-consistent, and auditable. We identify the Structure Gap as the deployment failure mode that makes Trace Integrity necessary: natural-language reasoning and free-form rationales do not reliably specify the operator-level programs required by real-world systems. We operationalize Trace Integrity with execution contracts, structured artifacts that bind user intent to schema elements, operator plans, assumptions, executable queries, verification status, and final-answer linkage. We also introduce CAIT (Correct Answer / Invalid Trace) Rate, which measures how often answer-only evaluation counts computationally unsupported outputs as successes. In an empirical demonstration on BIRD Mini-Dev, Direct SQL, Operation Summary + SQL, and Contract-First SQL achieve answer accuracies of 20%, 22%, and 24%, while their Trace Integrity Pass Rates are 39%, 43%, and 40% and their CAIT Rates remain high at 55%, 59.1%, and 45.8%, showing that answer accuracy, trace validity, and silent-failure risk are distinct evaluation signals. Real-world LLM data agents should, therefore, be evaluated not only by whether their outputs match a reference answer, but by whether those outputs are backed by auditable computation.

📄 PDF Abstract BibTeX arXiv:2608.26036

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Trusting What You Cannot See: Auditable Fine-Tuning and Inference for Proprietary AI

2026-03-08 · Heng Jin, Chaoyu Zhang, Hexuan Yu, Shanghao Shi 외 arxiv

Cloud-based infrastructure has become the dominant platform for deploying large models, particularly large language models (LLMs). Fine-tuning and inference are increasingly delegated to cloud providers for simplified de…

Data and Decision Traceability for SDA TAP Lab's Prototype Battle Management System

2025-02-13 · Latha Pratti, Samya Bagchi, Yasir Latif

Space Protocol is applying the principles derived from MITRE and NIST's Supply Chain Traceability: Manufacturing Meta-Framework (NIST IR 8536) to a complex multi party system to achieve introspection, auditing, and repla…

Management

Toward Auditable Neuro-Symbolic Reasoning in Pathology: SQL as an Explicit Trace of Evidence

2026-01-05 · Kewen Cao, Jianxu Chen, Yongbing Zhang, Ye Zhang 외 arxiv

Automated pathology image analysis is central to clinical diagnosis, but clinicians still ask which slide features drive a model's decision and why. Vision-language models can produce natural language explanations, but t…

Visual Question Answering

RewardHackingAgents: Benchmarking Evaluation Integrity for LLM ML-Engineering Agents

2026-03-11 · Yonas Atinafu, Robin Cohen arxiv

LLM agents increasingly perform end-to-end ML engineering tasks where success is judged by a single scalar test metric. This creates a structural vulnerability: an agent can increase the reported score by compromising th…

VeriGraph: Towards Verifiable Data-Analytic Agents

2026-06-15 · Jiajie Jin, Zhao Yang, Wenle Liao, Yuyang Hu 외 arxiv

LLM-based agents have demonstrated strong capabilities in data-intensive analytical tasks, yet their outputs are rarely verifiable: a reliance on linear text trajectories makes their reasoning difficult to audit. In part…