paper-with-me

Papers

Stop Shipping AI Agents on Faith: Capability Is Not Production Readiness

2026-07-30 · Fouad Bousetouane arxiv

AI agents are moving into production workflows where they retrieve information, call tools, maintain state, and act on behalf of users or organizations, but many release decisions still rely on capability signals, demos, or behavioral tests that do not show whether an agent is ready to operate under production constraints. Capability is therefore not production readiness. This paper introduces the ProofAgent Index (PAI), a governance readiness index for AI agents. PAI combines four dimensions of deployment evidence: Evaluation, Context, Compliance, and Governance. Evaluation measures observed behavior, Context measures the operating environment that shapes that behavior, Compliance measures alignment with applicable rules and controls, and Governance measures whether the organization can authorize, monitor, audit, and control the agent during operation. PAI is implemented inside ProofAgent Harness, an open source infrastructure for auditable AI agent evaluation and governance. Validation across two heavily regulated domains, healthcare and finance, shows that PAI carries held out readiness signal and separates higher risk from lower risk configurations. The results show that context engineering strongly changes reliability, capability improves behavior but does not determine readiness, and governance evidence must remain visible rather than averaged away. PAI reframes agent release from a faith based deployment decision into an auditable readiness decision.

📄 PDF Abstract BibTeX arXiv:2607.27677

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Learning Algorithms for Regenerative Stopping Problems with Applications to Shipping Consolidation in Logistics

2021-05-05 · Kishor Jothimurugan, Matthew Andrews, Jeongran Lee, Lorenzo Maggi

We study regenerative stopping problems in which the system starts anew whenever the controller decides to stop and the long-term average cost is to be minimized. Traditional model-based solutions involve estimating the …

Deep Reinforcement LearningImitation Learningreinforcement-learningReinforcement Learning (RL)

Beyond Code Generation: Reliability, Verification, and Cost Economics in the Agentic Software Development Lifecycle

2026-09-04 · Happy Bhati arxiv

AI coding systems are moving from autocomplete and chat toward agents that can inspect repositories, edit multiple files, run tools, write tests, open pull requests, and work for long periods with limited supervision. Th…

Code Generation

RouteCost: A Production-Inspired Multi-Stage Framework for Pre-Order Shipping Cost Estimation in E-Commerce

2026-06-24 · Xianling Zeng, Zihan Yu, Sichen Zhao, Yalun Qi 외 arxiv

Accurate pre-order shipping cost estimation is important in e-commerce because it affects price presentation, margin planning, and conversion. In practice, shipping cost is shaped not only by distance but also by destina…

The promising potential of vision language models for the generation of textual weather forecasts

2025-12-03 · Edward C. C. Steele, Dinesh Mane, Emilio Monti, Luis Orus 외 arxiv

Despite the promising capability of multimodal foundation models, their application to the generation of meteorological products and services remains nascent. To accelerate aspiration and adoption, we explore the novel u…

Understanding the Non-Convergence of Agricultural Futures via Stochastic Storage Costs and Timing Options

2017-04-11

This paper studies the market phenomenon of non-convergence between futures and spot prices in the grains market. We postulate that the positive basis observed at maturity stems from the futures holder's timing options t…