paper-with-me

홈 › Papers

SciPredict: Can LLMs Predict the Outcomes of Scientific Experiments in Natural Sciences?

2026-04-12 · Udari Madhushani Sehwag, Elaine Lau, Haniyeh Ehsani Oskouie, Shayan Shabihi, Erich Liang, Andrea Toledo, Guillermo Mangialardi, Sergio Fonrouge, Ed-Yeremai Hernandez Cardona, Paula Vergara, Utkarsh Tyagi, Chen Bo Calvin Zhang, Pavi Bhatter, Nicholas Johnson, Furong Huang, Ernesto Gabriel Hernandez Montoya, Bing Liu arxiv

Accelerating scientific discovery requires the identification of which experiments would yield the best outcomes before committing resources to costly physical validation. While existing benchmarks evaluate LLMs on scientific knowledge and reasoning, their ability to predict experimental outcomes - a task where AI could significantly exceed human capabilities - remains largely underexplored. We introduce SciPredict, a benchmark comprising 405 tasks derived from recent empirical studies in 33 specialized sub-fields of physics, biology, and chemistry. SciPredict addresses two critical questions: (a) can LLMs predict the outcome of scientific experiments with sufficient accuracy? and (b) can such predictions be reliably used in the scientific research process? Evaluations reveal fundamental limitations on both fronts. Model accuracies are 14-26% and human expert performance is $\approx$20%. Although some frontier models exceed human performance model accuracy is still far below what would enable reliable experimental guidance. Even within the limited performance, models fail to distinguish reliable predictions from unreliable ones, achieving only $\approx$20% accuracy regardless of their confidence or whether they judge outcomes as predictable without physical experimentation. Human experts, in contrast, demonstrate strong calibration: their accuracy increases from $\approx$5% to $\approx$80% as they deem outcomes more predictable without conducting the experiment. SciPredict establishes a rigorous framework demonstrating that superhuman performance in experimental science requires not just better predictions, but better awareness of prediction reliability. For reproducibility all our data and code are provided at https://github.com/scaleapi/scipredict

📄 PDF Abstract BibTeX arXiv:2604.10718

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Experiments or Outcomes? Probing Scientific Feasibility in Large Language Models

2026-04-20 · Seyedali Mohammadi, Manas Gaur, Francis Ferraro arxiv

Scientific feasibility assessment asks whether a claim is consistent with established knowledge and whether experimental evidence could support or refute it. We frame feasibility assessment as a diagnostic reasoning task…

Large language models surpass human experts in predicting neuroscience results

2024-03-04 · Xiaoliang Luo, Akilles Rechardt, Guangzhi Sun, Kevin K. Nejad 외

Scientific discoveries often hinge on synthesizing decades of research, a task that potentially outstrips human information processing capacities. Large language models (LLMs) offer a solution. LLMs trained on the vast s…

Reflections from the 2024 Large Language Model (LLM) Hackathon for Applications in Materials Science and Chemistry

2024-11-20 · Yoel Zimmermann, Adib Bazgir, Zartashia Afzal, Fariha Agbere 외

Here, we present the outcomes from the second Large Language Model (LLM) Hackathon for Applications in Materials Science and Chemistry, which engaged participants across global hybrid locations, resulting in 34 team subm…

Language ModelingLanguage ModellingLarge Language ModelProperty Prediction

Cooperative Profiles Predict Multi-Agent LLM Team Performance in AI for Science Workflows

2026-04-22 · Shivani Kumar, Adarsh Bharathwaj, David Jurgens arxiv

Multi-agent systems built from teams of large language models (LLMs) are increasingly deployed for collaborative scientific reasoning and problem-solving. These systems require agents to coordinate under shared constrain…

ExpVid: A Benchmark for Experiment Video Understanding & Reasoning

2025-10-13 · Yicheng Xu, Yue Wu, Jiashuo Yu, Ziang Yan 외 arxiv

Multimodal Large Language Models (MLLMs) hold promise for accelerating scientific discovery by interpreting complex experimental procedures. However, their true capabilities are poorly understood, as existing benchmarks …

Visual Grounding