paper-with-me

Papers

FrontierChallenge: Evaluating Scientific Workflow Completion

2026-08-25 · Liangcai Su, Zhaopeng Feng, Zhuo Chen, Zhen Zhang, Xiang Lin, Ruilin Li, Handuo Zhang, Ning Wang, Kailong Wen, Yueqi Guo, Feng Xing, Yiling Guo, Chenxiong Qian, Simon Shaolei Du, Lidong Bing, Xinyu Wang hf

Scientific agents increasingly analyze data, execute code, and produce research artifacts, yet most benchmarks emphasize final answers, isolated programs, or a single domain. We introduce FrontierChallenge, a cross-domain benchmark comprising 300 end-to-end scientific workflows. In this paper, we release and evaluate 97 of these tasks, spanning quantum chemistry, molecular dynamics, materials characterization, analytical chemistry, life science, and electrochemistry/environment. Each task provides fixed inputs and specifies a bundle of required scientific deliverables. We evaluate twelve frontier models with three agent scaffolds. Pass Rate measures the fraction of tasks satisfying the full-completion criterion, while Avg. Score captures partial progress. Each of the best-performing configurations completed only 20 of the 97 released tasks, yielding a Pass Rate of 20.6%. Partial progress translated especially poorly into complete delivery in analytical chemistry and electrochemistry/environment: Avg. Scores reached 87.6 and 94.9, but the highest Pass Rates were only 4% and 0%. Among non-passing Claude Code trajectories, 75.5% still ended with language claiming completion. Complementary HDS6 process scores correlate strongly with task outcomes, supporting FrontierChallenge as a benchmark of Heavy Duty Solver capabilities. These findings show that neither high partial scores nor confident claims of completion reliably indicate that a scientific task has been fully delivered, highlighting the need to evaluate end-to-end workflow execution and the completeness of scientific deliverables together.

📄 PDF Abstract BibTeX arXiv:2608.24979

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

SciExplore: Evaluating Autonomous Agents from Scientific Navigation to Information Integration

2026-07-23 · Yinhao Tang, Youqing Fang, Yanan Sun, Wenran Liu 외 arxiv

Scientific research involves complex information-seeking and reasoning workflows across heterogeneous sources. However, existing benchmarks primarily emphasize general-domain retrieval or static scientific question answe…

Question Answering

CheckSupport: A Local LLM-Powered Tool for Automated Manuscript Submission Checklist Selection and Completion

2026-05-10 · Satvik Tripathi, Don Enwerem, Kevin Song, Kristian Quevada 외 arxiv

Transparent and standardized reporting is essential for reproducible scientific research, yet adherence to reporting guidelines remains inconsistent because of the manual effort required to select and complete checklists…

Are LLMs Ready for Scientific Discovery? A Capability-Oriented Benchmark for AI Scientists

2026-07-13 · Chuhan Shi, Xiaoquan Ren, Sicheng Song, Haobo Li 외 hf

Existing benchmarks for scientific data analysis evaluate LLMs primarily on code execution or workflow completion, overlooking that scientific analysis serves to support distinct types of scientific claims: hypothesis ex…

Coding-agents can replicate scientific machine learning papers

2026-07-02 · Atharva Hans, Ilias Bilionis arxiv

Scientific machine learning papers typically make computational claims, e.g., that the relative mean square error is less than 5% or that the 95% predictive credible interval covers the test data. A coding agent can be p…

World of Workflows: A Benchmark for Bringing World Models to Enterprise Systems

2026-01-29 · Lakshya Gupta, Litao Li, Yizhe Liu, Sriram Ganapathi Subramanian 외 arxiv

Frontier large language models (LLMs) excel as autonomous agents in many domains, yet they remain untested in complex enterprise systems where hidden workflows create cascading effects across interconnected databases. Ex…