paper-with-me

Papers

Cheap Code, Costly Judgment: A Case Study on Governable Agentic Software Engineering

2026-07-01 · James C. Davis, Paschal C. Amusuo, Tanmay Singla, Berk Çakar, Kirsten A. Davis arxiv

Generative AI is shifting software engineering from a practice organized around scarce implementation effort toward one organized around abundant, low-cost code production. This shift changes the central engineering problem: not whether AI can generate useful code, but how engineers organize architectures, tools, evidence, and feedback loops so that AI-mediated development remains inspectable, correctable, and maintainable. We study this problem through a first-person case study: a 12-week development effort in which a single expert software engineer used frontier AI coding agents to build a document accessibility remediation system. The empirical record comprises 88 contemporaneous field notes, 420 KLOC of production code, and 1.16 MLOC of tests, lints, supporting documentation, and agent tooling. From this record, we develop a candidate middle-range theory of governance conversion, expressed as a process model explaining how high-velocity agentic implementation becomes governable. The model explains how agentic implementation velocity surfaces recurring structural failure classes, and how engineering judgment sustains velocity by converting those failures into durable governance mechanisms. In contrast to existing governance models that derive controls from known obligations, governance conversion explains how controls are discovered from failures that become visible only during agentic work. We use our model to make testable predictions and to describe implications for software engineering research and practice.

📄 PDF Abstract BibTeX arXiv:2607.01087

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Rich Insights from Cheap Signals: Efficient Evaluations via Tensor Factorization

2026-03-02 · Felipe Maia Polo, Aida Nematzadeh, Virginia Aglietti, Adam Fisch 외 arxiv

Moving beyond evaluations that collapse performance across heterogeneous prompts toward fine-grained evaluation at the prompt level, or within relatively homogeneous subsets, is necessary to diagnose generative models' s…

Limits to scalable evaluation at the frontier: LLM as Judge won't beat twice the data

2024-10-17 · Florian E. Dorner, Vivian Y. Nastl, Moritz Hardt

High quality annotations are increasingly a bottleneck in the explosively growing machine learning ecosystem. Scalable evaluation methods that avoid costly annotation have therefore become an important research ambition.…

Lowest-cost virus suppression

2021-02-09 · Jacob Janssen, Yaneer Bar-Yam

Analysis of policies for managing epidemics require simultaneously an economic and epidemiological perspective. We adopt a cost-of-policy framework to model both the virus spread and the cost of handling the pandemic. Be…

ACUTE-EVAL: Improved Dialogue Evaluation with Optimized Questions and Multi-turn Comparisons

2019-09-06 · Margaret Li, Jason Weston, Stephen Roller

While dialogue remains an important end-goal of natural language research, the difficulty of evaluation is an oft-quoted reason why it remains troublesome to make real progress towards its solution. Evaluation difficulti…

Dialogue Evaluation

pAI-Econ-claude: A Gated Human-in-the-Loop Multi-Agent Architecture for AI-Assisted Economic Theory Development

2026-07-23 · Chen Zhu, Xiaolu Wang, Weilong Zhang arxiv

In many social-science research tasks, such as economics, LLM-based agents must produce outputs for which no cheap, task-complete, machine-readable correctness signal exists. This creates a distinctive reliability proble…