A Critical Review of Causal Reasoning Benchmarks for Large Language Models
Numerous benchmarks aim to evaluate the capabilities of Large Language Models (LLMs) for causal inference and reasoning. However, many of them can likely be solved through the retrieval of domain knowledge, questioning whether they achieve their purpose. In this review, we present a comprehensive overview of LLM benchmarks for causality. We highlight how recent benchmarks move towards a more thorough definition of causal reasoning by incorporating interventional or counterfactual reasoning. We derive a set of criteria that a useful benchmark or set of benchmarks should aim to satisfy. We hope this work will pave the way towards a general framework for the assessment of causal understanding in LLMs and the design of novel benchmarks.
Code (0)
등록된 구현이 없습니다.
Tasks
Causal InferencecounterfactualCounterfactual ReasoningRetrievalMethods 이 논문이 사용한 방법론
Similar Papers 제목 키워드 기반
A Survey on Enhancing Causal Reasoning Ability of Large Language Models
Large language models (LLMs) have recently shown remarkable performance in language tasks and beyond. However, due to their limited inherent causal reasoning ability, LLMs still face challenges in handling tasks that req…
METER: Evaluating Multi-Level Contextual Causal Reasoning in Large Language Models
Contextual causal reasoning is a critical yet challenging capability for Large Language Models (LLMs). Existing benchmarks, however, often evaluate this skill in fragmented settings, failing to ensure context consistency…
BiasCause: Evaluate Socially Biased Causal Reasoning of Large Language Models
While large language models (LLMs) already play significant roles in society, research has shown that LLMs still generate content including social bias against certain sensitive groups. While existing benchmarks have eff…
BABE: Biology Arena BEnchmark
The rapid evolution of large language models (LLMs) has expanded their capabilities from basic dialogue to advanced scientific reasoning. However, existing benchmarks in biology often fail to assess a critical skill requ…
Unveiling Causal Reasoning in Large Language Models: Reality or Mirage?
Causal reasoning capability is critical in advancing large language models (LLMs) toward strong artificial intelligence. While versatile LLMs appear to have demonstrated capabilities in understanding contextual causality…
counterfactualGeneral Knowledge