paper-with-me

Papers

How Far Are We From True Auto-Research?

2026-05-18 · Zhengxin Zhang, Ning Wang, Sainyam Galhotra, Claire Cardie arxiv

Recent auto-research systems can produce complete papers, but feasibility is not the same as quality, and the field still lacks a systematic study of how good agent-generated papers actually are. We introduce ResearchArena, a minimal scaffold that lets off-the-shelf agents (Claude Code using Opus 4.6, Codex using GPT-5.4, and Kimi Code using K2.5) carry out the full research loop themselves (ideation, experimentation, paper writing, self-refinement) under only lightweight guidance. Across 13 computer science seeds and 3 trials per agent-domain pair, ResearchArena yields 117 agent-generated papers, each evaluated under three complementary lenses: a manuscript-only reviewer (SAR), an artifact-aware peer review (PR) in which agents inspect the workspace alongside the manuscript, and an human conducted meta-review. Under SAR alone the picture is optimistic: Claude Code obtains the highest score, outperforms Analemma's FARS, and matches the weighted-average human ICLR 2025 submission, suggesting that minimally scaffolded agents can produce papers that look competitive on manuscript-only review. Manual inspection, however, reveals this picture is overstated: SAR scores are poorly aligned with its actual acceptance decisions and reward plausible framing without verifying experimental substance. Under artifact-aware PR scores drop sharply, and manual auditing identifies experimental rigor as the major bottleneck, decomposing into three failure modes (fabricated results, underpowered experiments, and plan/execution mismatch) that are highly agent-dependent: Codex 5%/8% paper-vs-artifact mismatch / fabricated references versus Kimi Code 77%/72%, a $\sim$15$\times$ spread that tracks distinct research personas the agents develop. None of the 117 agent-generated papers reaches the acceptance bar of a top-tier venue. This suggests that we are still gapped from the true auto-research.

📄 PDF Abstract BibTeX arXiv:2605.19156

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Approximation Methods for Partially Observed Markov Decision Processes (POMDPs)

2021-08-31 · Caleb M. Bowyer

POMDPs are useful models for systems where the true underlying state is not known completely to an outside observer; the outside observer incompletely knows the true state of the system, and observes a noisy version of t…

Survey

Do Reservoir Computers Work Best at the Edge of Chaos?

2020-12-02 · Thomas L. Carroll

It has been demonstrated that cellular automata had the highest computational capacity at the edge of chaos, the parameter at which their behavior transitioned from ordered to chaotic. This same concept has been applied …

Using novel data and ensemble models to improve automated labeling of Sustainable Development Goals

2023-01-25 · Dirk U. Wulff, Dominik S. Meier, Rui Mata

A number of labeling systems based on text have been proposed to help monitor work on the United Nations (UN) Sustainable Development Goals (SDGs). Here, we present a systematic comparison of systems using a variety of t…

Specificity

Automatic True/False Question Generation for Educational Purpose

2022-07-01 · NAACL (BEA) 2022 7 · Bowei Zou, Pengfei Li, Liangming Pan, Ai Ti Aw

In field of teaching, true/false questioning is an important educational method for assessing students’ general understanding of learning materials. Manually creating such questions requires extensive human effort and ex…

Fact VerificationQuestion GenerationQuestion-GenerationReading Comprehension

Interpretation Gaps in LLM-Assisted Comprehension of Privacy Documents

2025-03-15 · Rinku Dewri

This article explores the gaps that can manifest when using a large language model (LLM) to obtain simplified interpretations of data practices from a complex privacy policy. We exemplify these gaps to showcase issues in…

Language ModelingLanguage ModellingLarge Language ModelManagement