paper-with-me

Papers

Evaluating Interventional Reasoning Capabilities of Large Language Models

2024-04-08 · Tejas Kasetty, Divyat Mahajan, Gintare Karolina Dziugaite, Alexandre Drouin, Dhanya Sridhar

Numerous decision-making tasks require estimating causal effects under interventions on different parts of a system. As practitioners consider using large language models (LLMs) to automate decisions, studying their causal reasoning capabilities becomes crucial. A recent line of work evaluates LLMs ability to retrieve commonsense causal facts, but these evaluations do not sufficiently assess how LLMs reason about interventions. Motivated by the role that interventions play in causal inference, in this paper, we conduct empirical analyses to evaluate whether LLMs can accurately update their knowledge of a data-generating process in response to an intervention. We create benchmarks that span diverse causal graphs (e.g., confounding, mediation) and variable types, and enable a study of intervention-based reasoning. These benchmarks allow us to isolate the ability of LLMs to accurately predict changes resulting from their ability to memorize facts or find other shortcuts. We evaluate six LLMs on the benchmarks, finding that GPT models show promising accuracy at predicting the intervention effects.

📄 PDF Abstract BibTeX arXiv:2404.05545

Code (0)

등록된 구현이 없습니다.

Tasks

Causal InferenceDecision Making

Methods 이 논문이 사용한 방법론

Refunds@Expedia|||How do I get a full refund from Expedia? “How do I get a full refund from Expedia? How do I get a full refund from Expedia? – Call ☎️ +1-(888) 829 (0881) or +1-805-330-4056 or +1-805-330-4056 for Quick Help &…
Attention 설명 없음
Cosine Annealing Cosine Annealing is a type of learning rate schedule that has the effect of starting with a large learning rate that is relatively rapidly decreased to a minimum value before…
Residual Connection 설명 없음
Adam 설명 없음
Weight Decay 설명 없음
Linear Warmup With Cosine Annealing Linear Warmup With Cosine Annealing is a learning rate schedule where we increase the learning rate linearly for $n$ updates and then anneal according to a cosine schedule…
Layer Normalization Unlike batch normalization, Layer Normalization directly estimates the normalization statistics from the summed inputs…

Similar Papers 제목 키워드 기반

Parallel Scaling Law: Unveiling Reasoning Generalization through A Cross-Linguistic Perspective

2025-10-02 · Wen Yang, Junhong Wu, Chong Li, Chengqing Zong 외 arxiv

Recent advancements in Reinforcement Post-Training (RPT) have significantly enhanced the capabilities of Large Reasoning Models (LRMs), sparking increased interest in the generalization of RL-based reasoning. While exist…

A Critical Review of Causal Reasoning Benchmarks for Large Language Models

2024-07-10 · Linying Yang, Vik Shirvaikar, Oscar Clivio, Fabian Falck

Numerous benchmarks aim to evaluate the capabilities of Large Language Models (LLMs) for causal inference and reasoning. However, many of them can likely be solved through the retrieval of domain knowledge, questioning w…

Causal InferencecounterfactualCounterfactual ReasoningRetrieval

Executable Counterfactuals: Improving LLMs' Causal Reasoning Through Code

2025-10-02 · Aniket Vashishtha, Qirun Dai, Hongyuan Mei, Amit Sharma 외 arxiv

Counterfactual reasoning, a hallmark of intelligence, consists of three steps: inferring latent variables from observations (abduction), constructing alternatives (interventions), and predicting their outcomes (predictio…

Reinforcement Learning

CLadder: Assessing Causal Reasoning in Language Models

2023-12-07 · NeurIPS 2023 11 · Zhijing Jin, Yuen Chen, Felix Leeb, Luigi Gresele 외

The ability to perform causal reasoning is widely considered a core feature of intelligence. In this work, we investigate whether large language models (LLMs) can coherently reason about causality. Much of the existing w…

Causal InferenceCommonsense Causal Reasoningcounterfactual

CARL-GT: Evaluating Causal Reasoning Capabilities of Large Language Models

2024-12-23 · Ruibo Tu, Hedvig Kjellström, Gustav Eje Henter, Cheng Zhang

Causal reasoning capabilities are essential for large language models (LLMs) in a wide range of applications, such as education and healthcare. But there is still a lack of benchmarks for a better understanding of such c…

Decision MakingMathZero-Shot Learning