paper-with-me

홈 › Papers

SciCode: A Research Coding Benchmark Curated by Scientists

2024-07-18 · Minyang Tian, Luyu Gao, Shizhuo Dylan Zhang, Xinan Chen, Cunwei Fan, Xuefei Guo, Roland Haas, Pan Ji, Kittithat Krongchon, Yao Li, Shengyan Liu, Di Luo, Yutao Ma, Hao Tong, Kha Trinh, Chenyu Tian, Zihan Wang, Bohao Wu, Yanyu Xiong, Shengzhu Yin, Minhui Zhu, Kilian Lieret, Yanxin Lu, Genglin Liu, Yufeng Du, Tianhua Tao, Ofir Press, Jamie Callan, Eliu Huerta, Hao Peng

Since language models (LMs) now outperform average humans on many challenging tasks, it has become increasingly difficult to develop challenging, high-quality, and realistic evaluations. We address this issue by examining LMs' capabilities to generate code for solving real scientific research problems. Incorporating input from scientists and AI researchers in 16 diverse natural science sub-fields, including mathematics, physics, chemistry, biology, and materials science, we created a scientist-curated coding benchmark, SciCode. The problems in SciCode naturally factorize into multiple subproblems, each involving knowledge recall, reasoning, and code synthesis. In total, SciCode contains 338 subproblems decomposed from 80 challenging main problems. It offers optional descriptions specifying useful scientific background information and scientist-annotated gold-standard solutions and test cases for evaluation. Claude3.5-Sonnet, the best-performing model among those tested, can solve only 4.6% of the problems in the most realistic setting. We believe that SciCode demonstrates both contemporary LMs' progress towards becoming helpful scientific assistants and sheds light on the development and evaluation of scientific AI in the future.

📄 PDF Abstract BibTeX arXiv:2407.13168

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Scientific Code Search at Scale: A Multi-Domain Dataset and Benchmark

2026-07-03 · Nishan Pantha, Pranath Reddy Kumbam, Sajil Awale, Pushwitha Krishnappa 외 arxiv

Scientists increasingly rely on open-source tools to support their research workflows, yet discovering relevant software among over 600 million GitHub repositories remains challenging. Existing code search benchmarks foc…

Information RetrievalCode Search

scicode-lint: Detecting Methodology Bugs in Scientific Python Code with LLM-Generated Patterns

2026-03-18 · Sergey V. Samsonau arxiv

Methodology bugs in scientific Python code produce plausible but incorrect results that traditional linters and static analysis tools cannot detect. Several research groups have built ML-specific linters, demonstrating t…

Toward a Team of AI-made Scientists for Scientific Discovery from Gene Expression Data

2024-02-15 · Haoyang Liu, Yijiang Li, Jinglin Jian, Yuxuan Cheng 외

Machine learning has emerged as a powerful tool for scientific discovery, enabling researchers to extract meaningful insights from complex datasets. For instance, it has facilitated the identification of disease-predicti…

Language ModelingLanguage ModellingLarge Language Modelscientific discovery

Harnessing AtomisticSkills for Agentic Atomistic Research

2026-05-18 · Bowen Deng, Bohan Li, Matthew Cox, Hoje Chun 외 arxiv

Computational materials science and chemistry span vast knowledge domains and fractured software ecosystems. Although large language models (LLMs) have demonstrated research capabilities, scaling monolithic agents to man…

Drug Discovery

AutoResearch: An Execution-Grounded Multi-Agent Framework for Reliable Research Workflow Automation

2026-05-04 · Rajesh Kumar, Waqar Ali, Junaid Ahmed, Abdullah Aman Khan 외 arxiv

Automated research agents increasingly generate code, retrieve literature, and draft scientific artifacts, but they often fail to verify whether generated experiments execute correctly or whether cited sources support ge…

Code Repair