paper-with-me

홈 › Papers

Training AI Scientists to Replicate Research

2026-08-13 · Damon Falck, Samer Sabri, Anja Surina, Thom Foster, Anya Sims, Sam Devlin, Dylan Rogers, Tantum Collins, Kaloyan Aleksiev, Louis Kirsch, Edward Hughes arxiv

The replicability of papers is a cornerstone of scientific knowledge, ensuring the reliability of existing results and providing a base for further experiments. The act of replication typically illuminates details that were previously underspecified, and thus requires similar hypothesis-driven exploration to open-ended research. In this work, we develop Replica, a scalable task space for paper replication. To provide reward signal, we introduce an auto-generated rubric-based judge that has low noise and agrees with human assessment of replication quality. We post-train Faraday, a 27B-parameter "AI Scientist" agent that leverages coding agents as tools, surpassing the performance of Claude Opus 4.8 and GPT-5.5 on held-out replication tasks. Qualitative analysis of individual rollouts reveals that Faraday adopts a more scientifically-principled approach. We believe that our results provide a stepping stone towards AI agents capable of long-horizon scientific innovation without requiring complex harnesses.

📄 PDF Abstract BibTeX arXiv:2608.13331

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Toward Reliable Ad-hoc Scientific Information Extraction: A Case Study on Two Materials Datasets

2024-06-08 · Satanu Ghosh, Neal R. Brodnik, Carolina Frey, Collin Holgate 외

We explore the ability of GPT-4 to perform ad-hoc schema based information extraction from scientific literature. We assess specifically whether it can, with a basic prompting approach, replicate two existing material sc…

Toward a Team of AI-made Scientists for Scientific Discovery from Gene Expression Data

2024-02-15 · Haoyang Liu, Yijiang Li, Jinglin Jian, Yuxuan Cheng 외

Machine learning has emerged as a powerful tool for scientific discovery, enabling researchers to extract meaningful insights from complex datasets. For instance, it has facilitated the identification of disease-predicti…

Language ModelingLanguage ModellingLarge Language Modelscientific discovery

Selecting Language Models for Social Science: Start Small, Start Open, and Validate

2026-01-16 · Dustin S. Stoltz, Marshall A. Taylor, Sanuj Kumar arxiv

Currently, there are thousands of large pretrained language models (LLMs) available to social scientists. How do we select among them? Using validity, reliability, reproducibility, and replicability as guides, we explore…

Can GPT-4 Replicate Empirical Software Engineering Research?

2023-10-03 · Jenny T. Liang, Carmen Badea, Christian Bird, Robert DeLine 외

Empirical software engineering research on production systems has brought forth a better understanding of the software engineering process for practitioners and researchers alike. However, only a small subset of producti…

The structure of behavioral data

2020-12-23 · Aurélien Defossez, Morteza Ansarinia, Brice Clocher, Emmanuel Schmück 외

For more than a century, scientists have been collecting behavioral data--an increasing fraction of which is now being publicly shared so other researchers can reuse them to replicate, integrate or extend past results. A…