paper-with-me

홈 › Papers

When simulations look right but causal effects go wrong: Large language models as behavioral simulators

2026-04-02 · Zonghan Li, Feng Ji arxiv

Behavioral simulation is increasingly used to anticipate responses to interventions. Large language models (LLMs) enable researchers to specify population characteristics and intervention context in natural language, but it remains unclear to what extent LLMs can use these inputs to infer intervention effects. We evaluated three LLMs on 11 climate-psychology interventions using a dataset of 59,508 participants from 62 countries, and replicated the main analysis in two additional datasets (12 and 27 countries). LLMs reproduced observed patterns in attitudinal outcomes (e.g., climate beliefs and policy support) reasonably well, and prompting refinements improved this descriptive fit. However, descriptive fit did not reliably translate into causal fidelity (i.e., accurate estimates of intervention effects), and these two dimensions of accuracy followed different error structures. This descriptive-causal divergence held across the three datasets, but varied across intervention logics, with larger errors for interventions that depended on evoking internal experience than on directly conveying reasons or social cues. It was more pronounced for behavioral outcomes, where LLMs imposed stronger attitude-behavior coupling than in human data. Countries and population groups appearing well captured descriptively were not necessarily those with lower causal errors. Relying on descriptive fit alone may therefore create unwarranted confidence in simulation results, misleading conclusions about intervention effects and masking population disparities that matter for fairness.

📄 PDF Abstract BibTeX arXiv:2604.02458

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Incorporating structural uncertainty in causal decision making

2025-07-31 · Maurits Kaptein arxiv

Practitioners making decisions based on causal effects typically ignore structural uncertainty. We analyze when this uncertainty is consequential enough to warrant methodological solutions (Bayesian model averaging over …

Causal InferenceDecision Making

This human study did not involve human subjects: Validating LLM simulations as behavioral evidence

2026-02-17 · Jessica Hullman, David Broska, Huaman Sun, Aaron Shaw arxiv

A growing literature uses large language models (LLMs) as synthetic participants to generate cost-effective and nearly instantaneous responses in social science experiments. However, there is limited guidance on when suc…

Prompt Engineering

Estimating Spillover Effects in the Presence of Isolated Nodes

2024-12-08 · Bora Kim

In estimating spillover effects under network interference, practitioners often use linear regression with either the number or fraction of treated neighbors as regressors. An often overlooked fact is that the latter is …

regression

Using Time Structure to Estimate Causal Effects

2025-04-15 · Tom Hochsprung, Jakob Runge, Andreas Gerhardus

There exist several approaches for estimating causal effects in time series when latent confounding is present. Many of these approaches rely on additional auxiliary observed variables or time series such as instruments,…

Time Series

Estimating heterogeneous treatment effects with right-censored data via causal survival forests

2020-01-27 · Yifan Cui, Michael R. Kosorok, Erik Sverdrup, Stefan Wager 외

Forest-based methods have recently gained in popularity for non-parametric treatment effect estimation. Building on this line of work, we introduce causal survival forests, which can be used to estimate heterogeneous tre…