paper-with-me

홈 › Papers

Toward Autonomous Long-Horizon Engineering for ML Research

2026-04-14 · Guoxin Chen, Jie Chen, Lei Chen, Jiale Zhao, Fanzhe Meng, Wayne Xin Zhao, Ruihua Song, Cheng Chen, Ji-Rong Wen, Kai Jia arxiv

Agentic systems increasingly automate pieces of AI research. Yet turning underspecified research objectives into runnable, experimentally validated ML systems remains a central bottleneck. We study this operational setting as \emph{long-horizon ML research engineering}: converting a research specification into a runnable ML system through repeated implementation, experimentation, and refinement. The central challenge is to sustain cumulative project progress across heterogeneous stages under delayed, confounded feedback. We introduce AiScientist, a multi-agent system built around thin control over thick state: a lightweight hierarchical research team coordinates through a File-as-Bus workspace that preserves decision-relevant artifacts across roles and invocations. On PaperBench, AiScientist improves over the strongest matched baselines by 9.92 and 11.15 points with Gemini-3-Flash and GLM-5, respectively. On MLE-Bench Lite, it reaches 81.82 Any Medal\% under both backbones, improving over the strongest matched baselines by 4.55 and 16.67 points, and exceeding a Codex/GPT-5.5 xhigh frontier harness reference by 13.64 Any Medal points. Ablations and process analyses show that durable project state is central to later-round refinement: removing File-as-Bus lowers PaperBench score by 6.41 points and MLE-Bench Lite Any Medal\% by 31.82 points. These results suggest that long-horizon AI research is not only a problem of stronger local reasoning, but a systems problem of maintaining cumulative, inspectable project progress.

📄 PDF Abstract BibTeX arXiv:2604.13018

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Toward Ultra-Long-Horizon Agentic Science: Cognitive Accumulation for Machine Learning Engineering

2026-01-15 · Xinyu Zhu, Yuzhu Cai, Zexi Liu, Bingyang Zheng 외 arxiv

The advancement of artificial intelligence toward agentic science is currently bottlenecked by the challenge of ultra-long-horizon autonomy, the ability to sustain strategic coherence and iterative correction over experi…

AutoLab: Can Frontier Models Solve Long-Horizon Auto Research and Engineering Tasks?

2026-06-03 · Zhangchen Xu, Junda Chen, Yue Huang, Dongfu Jiang 외 arxiv

Scientific and engineering progress is fundamentally a long-horizon iterative process: proposing changes, running experiments, measuring outcomes, and continuously refining artifacts. Yet existing benchmarks for frontier…

Φ-Bench: Can Large Language Models Engineer the Infrastructure That Powers Them?

2026-09-09 · Leilei Ding, Shumin Wang, Yuting Huang, Fanqi Wan 외 arxiv

Large language models (LLMs) have demonstrated remarkable capabilities in reasoning and code generation, raising the prospect that they could assist in developing and optimizing the very infrastructure that powers them. …

Code Generation

Beyond Final Scores: A Systematic Evaluation of Agents for Long-Horizon AI Research and Development

2026-08-13 · Yiwei Li, Wanli Yang, Hexiang Tan, Xiangzhou Huang 외 arxiv

Autonomous agents are increasingly capable of improving models, systems, and other technical artifacts through long-horizon experimentation. To understand the current state of this capability, however, evaluation must go…

Omni-SimpleMem: Autoresearch-Guided Discovery of Lifelong Multimodal Agent Memory

2026-04-01 · Jiaqi Liu, Zipeng Ling, Shi Qiu, Yanqing Liu 외 arxiv

AI agents increasingly operate over extended time horizons, yet their ability to retain, organize, and recall multimodal experiences remains a critical bottleneck. Building effective lifelong memory requires navigating a…

Prompt Engineering