paper-with-me

Papers

A Critical Review of Causal Reasoning Benchmarks for Large Language Models

2024-07-10 · Linying Yang, Vik Shirvaikar, Oscar Clivio, Fabian Falck

Numerous benchmarks aim to evaluate the capabilities of Large Language Models (LLMs) for causal inference and reasoning. However, many of them can likely be solved through the retrieval of domain knowledge, questioning whether they achieve their purpose. In this review, we present a comprehensive overview of LLM benchmarks for causality. We highlight how recent benchmarks move towards a more thorough definition of causal reasoning by incorporating interventional or counterfactual reasoning. We derive a set of criteria that a useful benchmark or set of benchmarks should aim to satisfy. We hope this work will pave the way towards a general framework for the assessment of causal understanding in LLMs and the design of novel benchmarks.

📄 PDF Abstract BibTeX arXiv:2407.08029

Code (0)

등록된 구현이 없습니다.

Tasks

Causal InferencecounterfactualCounterfactual ReasoningRetrieval

Methods 이 논문이 사용한 방법론

SET Dynamic Sparse Training method where weight mask is updated randomly periodically
Causal inference Causal inference is the process of drawing a conclusion about a causal connection based on the conditions of the occurrence of an effect. The main difference between causal…

Similar Papers 제목 키워드 기반

A Survey on Enhancing Causal Reasoning Ability of Large Language Models

2025-03-12 · Xin Li, Zhuo Cai, Shoujin Wang, Kun Yu 외

Large language models (LLMs) have recently shown remarkable performance in language tasks and beyond. However, due to their limited inherent causal reasoning ability, LLMs still face challenges in handling tasks that req…

METER: Evaluating Multi-Level Contextual Causal Reasoning in Large Language Models

2026-04-13 · Pengfeng Li, Chen Huang, Chaoqun Hao, Hongyao Chen 외 arxiv

Contextual causal reasoning is a critical yet challenging capability for Large Language Models (LLMs). Existing benchmarks, however, often evaluate this skill in fragmented settings, failing to ensure context consistency…

BiasCause: Evaluate Socially Biased Causal Reasoning of Large Language Models

2025-04-08 · Tian Xie, Tongxin Yin, Vaishakh Keshava, Xueru Zhang 외

While large language models (LLMs) already play significant roles in society, research has shown that LLMs still generate content including social bias against certain sensitive groups. While existing benchmarks have eff…

BABE: Biology Arena BEnchmark

2026-02-05 · Junting Zhou, Jin Chen, Linfeng Hao, Denghui Cao 외 arxiv

The rapid evolution of large language models (LLMs) has expanded their capabilities from basic dialogue to advanced scientific reasoning. However, existing benchmarks in biology often fail to assess a critical skill requ…

Unveiling Causal Reasoning in Large Language Models: Reality or Mirage?

2025-06-26 · Haoang Chi, He Li, Wenjing Yang, Feng Liu 외

Causal reasoning capability is critical in advancing large language models (LLMs) toward strong artificial intelligence. While versatile LLMs appear to have demonstrated capabilities in understanding contextual causality…

counterfactualGeneral Knowledge