paper-with-me

홈 › Papers

ReplicationBench: Can AI Agents Replicate Astrophysics Research Papers?

2025-10-28 · Christine Ye, Sihan Yuan, Suchetha Cooray, Steven Dillmann, Ian L. V. Roque, Dalya Baron, Philipp Frank, Sergio Martin-Alvarez, Nolan Koblischke, Frank J Qu, Diyi Yang, Risa Wechsler, Ioana Ciuca arxiv

Frontier AI agents show increasing promise as scientific research assistants, and may eventually be useful for extended, open-ended research workflows. However, in order to use agents for novel research, we must first assess the underlying faithfulness and correctness of their work. To evaluate agents as research assistants, we introduce ReplicationBench, an evaluation framework that tests whether agents can replicate entire research papers drawn from the astrophysics literature. Astrophysics, where research relies heavily on archival data and computational study while requiring little real-world experimentation, is a particularly useful testbed for AI agents in scientific research. We split each paper into tasks which require agents to replicate the paper's core contributions, including the experimental setup, derivations, data analysis, and codebase. Each task is co-developed with the original paper authors and targets a key scientific result, enabling objective evaluation of both faithfulness (adherence to original methods) and correctness (technical accuracy of results). ReplicationBench is extremely challenging for current frontier language models: even the best-performing language models score under 20%. We analyze ReplicationBench trajectories in collaboration with domain experts and find a rich, diverse set of failure modes for agents in scientific research. ReplicationBench establishes the first benchmark of paper-scale, expert-validated astrophysics research tasks, reveals insights about agent performance generalizable to other domains of data-driven science, and provides a scalable framework for measuring AI agents' reliability in scientific research.

📄 PDF Abstract BibTeX arXiv:2510.24591

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

VERITAS: Towards a General-Purpose Replication Tool for Scientific Research

2026-07-03 · Haokun Liu, Filbert Aurelian Tjiaranata, Chenhao Tan arxiv

AI tools are accelerating scientific publication while the systems that review it struggle to keep up, and independent verification of published research has become both harder and more important. As manual replication i…

Improving astroBERT using Semantic Textual Similarity

2022-11-29 · Felix Grezes, Thomas Allen, Sergi Blanco-Cuaresma, Alberto Accomazzi 외

The NASA Astrophysics Data System (ADS) is an essential tool for researchers that allows them to explore the astronomy and astrophysics scientific literature, but it has yet to exploit recent advances in natural language…

AstronomyLanguage ModelingLanguage ModellingSemantic Textual Similarity

PaperBench: Evaluating AI's Ability to Replicate AI Research

2025-04-02 · Giulio Starace, Oliver Jaffe, Dane Sherburn, James Aung 외

We introduce PaperBench, a benchmark evaluating the ability of AI agents to replicate state-of-the-art AI research. Agents must replicate 20 ICML 2024 Spotlight and Oral papers from scratch, including understanding paper…

AI's Capability in Assisting Scientific Research in Physics, Astrophysics, and Cosmology I: Literature Review

2026-07-28 · Anamaria Hell, Kateryna Vovk, Veena Krishnaraj, Jia Liu 외 arxiv

We investigate how well large language models (LLMs) can assist with literature reviews for scientific research. We perform a controlled study of eight expert-conceived research projects across the areas of physics, astr…

ReplicatorBench: Benchmarking LLM Agents for Replicability in Social and Behavioral Sciences

2026-02-11 · Bang Nguyen, Dominik Soós, Qian Ma, Rochana R. Obadage 외 arxiv

The literature has witnessed an emerging interest in AI agents for automated assessment of scientific papers. Existing benchmarks focus primarily on the computational aspect of this task, testing agents' ability to repro…