paper-with-me

홈 › Papers

Read the Paper, Write the Code: Agentic Reproduction of Social-Science Results

2026-04-23 · Benjamin Kohler, David Zollikofer, Johanna Einsiedler, Alexander Hoyle, Elliott Ash arxiv

Recent work has used LLM agents to reproduce empirical social science results with access to both the data and code. We broaden this scope by asking: Can they reproduce results given only a paper's methods description and original data? We develop an agentic reproduction system that extracts structured methods descriptions from papers, runs reimplementations under strict information isolation -- agents never see the original code, results, or paper -- and enables deterministic, cell-level comparison of reproduced outputs to the original results. An error attribution step traces discrepancies through the system chain to identify root causes. Evaluating four agent scaffolds and four LLMs on 48 papers with human-verified reproducibility, we find that agents can largely recover published results, but performance varies substantially between models, scaffolds, and papers. Root cause analysis reveals that failures stem both from agent errors and from underspecification in the papers themselves.

📄 PDF Abstract BibTeX arXiv:2604.21965

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

From paper to benchmark: agentic, framework-based reproduction of under-specified methods in machine health intelligence

2026-05-27 · Raffael Theiler, Ludovico Comito, David Leko, Leandro Von Krannichfeldt 외 arxiv

Industrial Prognostics and Health Management (PHM) provides a representative case study for a broader challenge in applied machine learning: translating published papers into executable, benchmark-ready implementations. …

TDFlow: Agentic Workflows for Test Driven Development

2025-10-27 · Kevin Han, Siddharth Maddikayala, Tim Knappe, Om Patel 외 arxiv

We introduce TDFlow, a novel test-driven agentic workflow that frames repository-scale software engineering as a test-resolution task, specifically designed to solve human-written tests. Given a set of tests, TDFlow repe…

Program Repair

Beyond Human-Readable: Rethinking Software Engineering Conventions for the Agentic Development Era

2026-04-08 · Dmytro Ustynov arxiv

For six decades, software engineering principles have been optimized for a single consumer: the human developer. The rise of agentic AI development, where LLM-based agents autonomously read, write, navigate, and debug co…

Bad Memory: Evaluating Prompt Injection Risks from Memory in Agentic Systems

2026-07-16 · Soham Gadgil, David Alexander, Sai Sunku, Franziska Roesner arxiv

A growing class of agentic systems maintain persistent state across sessions through memory files, behavioral preferences, and knowledge bases. While this makes agents more useful and self-improving, it also creates a ne…

Remember Your Trace: Memory-Guided Long-Horizon Agentic Framework for Consistent and Hierarchical Repository-Level Code Documentation

2026-05-14 · Suyoung Bae, Jaehoon Lee, Changkyu Choi, YunSeok Choi 외 arxiv

Automated code documentation is essential for modern software development, providing the contextual grounding that both human developers and coding agents rely on to navigate large codebases. Existing repository-level ap…