paper-with-me

홈 › Papers

Are LLMs Ready for Scientific Discovery? A Capability-Oriented Benchmark for AI Scientists

2026-07-13 · Chuhan Shi, Xiaoquan Ren, Sicheng Song, Haobo Li, Rui Sheng, Yushi Sun hf

Existing benchmarks for scientific data analysis evaluate LLMs primarily on code execution or workflow completion, overlooking that scientific analysis serves to support distinct types of scientific claims: hypothesis exploration, statistical inference, mechanistic explanation, each with different assumptions and validity criteria. We introduce SDABench, a benchmark that reorganizes evaluation around six capabilities (descriptive, exploratory, inferential, predictive, causal, and mechanistic) across five domains (Biology, Chemistry, Environment, Geography, Physics). SDABench comprises 527 real-data instances (SDA-Real) and 6000 synthetic instances (SDA-Synth), each in both multiple-choice and open-ended formats, constructed through an automated pipeline. Evaluating 15 representative LLMs, we find that models handle descriptive analysis well but degrade sharply on tasks requiring assumption selection, latent-process modeling, or mechanistic reasoning. SDABench further provides a five-stage error analysis framework that locates where LLMs fail: more advanced models more reliably identify the relevant scope and variables, but still struggle to select appropriate analytical procedures, model variable relationships, and draw valid conclusions.

📄 PDF Abstract BibTeX arXiv:2607.11079

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Evolving Roles of LLMs in Scientific Innovation: Assistant, Collaborator, Scientist, and Evaluator

2025-07-16 · Haoxuan Zhang, Ruochi Li, Yang Zhang, Ting Xiao 외 arxiv

Large language models (LLMs) are increasingly used in scientific research and discovery, supporting tasks ranging from literature retrieval and synthesis to hypothesis generation, autonomous experimentation, and research…

From AI for Science to Agentic Science: A Survey on Autonomous Scientific Discovery

2025-08-18 · Jiaqi Wei, Yuejin Yang, Xiang Zhang, Yuhan Chen 외 arxiv

Artificial intelligence (AI) is reshaping scientific discovery, evolving from specialized computational tools into autonomous research partners. We position Agentic Science as a pivotal stage within the broader AI for Sc…

Evaluating Large Language Models in Scientific Discovery

2025-12-17 · Zhangde Song, Jieyu Lu, Yuanqi Du, Botao Yu 외 arxiv

Large language models (LLMs) are increasingly applied to scientific research, yet prevailing science benchmarks probe decontextualized knowledge and overlook the iterative reasoning, hypothesis generation, and observatio…

SciHorizon: Benchmarking AI-for-Science Readiness from Scientific Data to Large Language Models

2025-03-12 · Chuan Qin, Xin Chen, Chengrui Wang, Pengmin Wu 외

In recent years, the rapid advancement of Artificial Intelligence (AI) technologies, particularly Large Language Models (LLMs), has revolutionized the paradigm of scientific discovery, establishing AI-for-Science (AI4Sci…

BenchmarkingFairnessscientific discovery

Revisiting Gene Ontology Knowledge Discovery with Hierarchical Feature Selection and Virtual Study Group of AI Agents

2026-03-20 · Cen Wan, Alex A. Freitas arxiv

Large language models have achieved great success in multiple challenging tasks, and their capacity can be further boosted by the emerging agentic AI techniques. This new computing paradigm has already started revolution…