paper-with-me

홈 › Papers

Realizing LLMs' Causal Potential Requires Science-Grounded, Novel Benchmarks

2025-10-18 · Ashutosh Srivastava, Lokesh Nagalapatti, Gautam Jajoo, Aniket Vashishtha, Parameswari Krishnamurthy, Amit Sharma arxiv

Recent claims of strong performance by Large Language Models (LLMs) on causal discovery are undermined by a key flaw: many evaluations rely on benchmarks likely included in pretraining corpora. Thus, apparent success suggests that LLM-only methods, which ignore observational data, outperform classical statistical approaches. We challenge this narrative by asking: Do LLMs truly reason about causal structure, and how can we measure it without memorization concerns? Can they be trusted for real-world scientific discovery? We argue that realizing LLMs' potential for causal analysis requires two shifts: (P.1) developing robust evaluation protocols based on recent scientific studies to guard against dataset leakage, and (P.2) designing hybrid methods that combine LLM-derived knowledge with data-driven statistics. To address P.1, we encourage evaluating discovery methods on novel, real-world scientific studies. We outline a practical recipe for extracting causal graphs from recent publications released after an LLM's training cutoff, ensuring relevance and preventing memorization while capturing both established and novel relations. Compared to benchmarks like BNLearn, where LLMs achieve near-perfect accuracy, they perform far worse on our curated graphs, underscoring the need for statistical grounding. Supporting P.2, we show that using LLM predictions as priors for the classical PC algorithm significantly improves accuracy over both LLM-only and purely statistical methods. We call on the community to adopt science-grounded, leakage-resistant benchmarks and invest in hybrid causal discovery methods suited to real-world inquiry.

📄 PDF Abstract BibTeX arXiv:2510.16530

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

InterveneBench: Benchmarking LLMs for Intervention Reasoning and Causal Study Design in Real Social Systems

2026-03-16 · Shaojie Shi, Zhengyu Shi, Lingran Zheng, Xinyu Su 외 arxiv

Causal inference in social science relies on end-to-end, intervention-centered research-design reasoning grounded in real-world policy interventions, but current benchmarks fail to evaluate this capability of large langu…

Causal Inference

Born-Qualified: An Autonomous Framework for Deploying Advanced Energy and Electronic Materials

2026-05-01 · Steven R. Spurgeon, Milad Abolhasani, Frederick Baddour, Ryan B. Comes 외 arxiv

Autonomous science is transforming how we discover materials and chemical systems for advanced energy technologies. However, many initially promising systems never reach deployment. This "valley of death" stems from opti…

ParKCa: Causal Inference with Partially Known Causes

2020-03-17 · Raquel Aoki, Martin Ester

Methods for causal inference from observational data are an alternative for scenarios where collecting counterfactual data or realizing a randomized experiment is not possible. Adopting a stacking approach, our proposed …

Causal Inferencecounterfactual

NeuroAI and Beyond: Bridging Between Advances in Neuroscience and ArtificialIntelligence

2026-04-19 · Anthony Zador, Jean-Marc Fellous, Terrence Sejnowski, Gina Adam 외 arxiv

Neuroscience and Artificial Intelligence (AI) have made impressive progress in recent years but remain only loosely interconnected. Based on a workshop convened by the National Science Foundation in August 2025, we ident…

Do LLMs Act as Repositories of Causal Knowledge?

2024-12-14 · Nick Huntington-Klein, Eleanor J. Murray

Large language models (LLMs) offer the potential to automate a large number of tasks that previously have not been possible to automate, including some in science. There is considerable interest in whether LLMs can autom…

Causal InferenceMultiple-choice