paper-with-me

홈 › Papers

Generative Synthetic Data for Causal Inference: Pitfalls, Remedies, and Opportunities

2026-04-26 · Yichen Xu arxiv

Synthetic tabular data are often evaluated by distributional similarity, privacy distance, or train-on-synthetic-test-on-real predictive performance, but these criteria do not ensure validity for causal inference. We show that fully generative tabular synthesizers, including GAN- and LLM-based models, can preserve predictive utility while distorting average treatment effect (ATE) estimates. The failure is structural: ATE preservation requires both a realistic covariate law and an accurate treatment-effect contrast, whereas prediction loss penalizes treatment-effect error only through an overlap-weighted term. Thus, under imbalance or limited overlap, a generator may reproduce dominant observed outcomes while underlearning intervention-relevant contrasts. We formalize this mismatch through sensitivity and loss-decomposition results. Motivated by this causal analysis and intuition, we propose a hybrid synthetic-data framework for causal inference that generates covariates while modeling treatment and outcome mechanisms separately. We evaluate the framework in three settings: ATE preservation under fully generative versus hybrid synthesis, augmentation for practical positivity problems, and diagnostic simulation engines for comparing OR, IPW, AIPW, and TMLE before real-data analysis. We also stress-test the hybrid construction across settings that vary overlap, covariate dimension, seed sample size, and treatment-effect complexity, including a logistic outcome-model misspecification check. Across controlled simulation experiments, hybrid synthesis improves causal fidelity relative to fully generative baselines; the ACTG application shows improved predictive fidelity and potential for finite-sample estimator benchmarking. LLM-based hybrid synthesis is often more faithful than CTGAN in settings where causal fidelity can be assessed.

📄 PDF Abstract BibTeX arXiv:2604.23904

Code (0)

등록된 구현이 없습니다.

Tasks

Causal Inference

Similar Papers 제목 키워드 기반

High Dimensional Causal Inference with Variational Backdoor Adjustment

2023-10-09 · Daniel Israel, Aditya Grover, Guy Van Den Broeck

Backdoor adjustment is a technique in causal inference for estimating interventional quantities from purely observational data. For example, in medical settings, backdoor adjustment can be used to control for confounding…

Causal InferenceVariational Inference

Ice Cream Doesn't Cause Drowning: Benchmarking LLMs Against Statistical Pitfalls in Causal Inference

2025-05-19 · Jin Du, Li Chen, Xun Xian, an Luo 외

Reliable causal inference is essential for making decisions in high-stakes areas like medicine, economics, and public policy. However, it remains unclear whether large language models (LLMs) can handle rigorous and trust…

BenchmarkingCausal InferenceSelection bias

Improving Generative Methods for Causal Evaluation via Simulation-Based Inference

2025-09-02 · Pracheta Amaranath, Vinitra Muralikrishnan, Amit Sharma, David Jensen arxiv

Generating synthetic datasets that accurately reflect real-world observational data is critical for evaluating causal estimators, but it remains a challenging task. Existing generative methods offer a solution by produci…

Harnessing Synthetic Data from Generative AI for Statistical Inference

2026-03-05 · Ahmad Abdel-Azim, Ruoyu Wang, Xihong Lin arxiv

The emergence of generative AI models has dramatically expanded the availability and use of synthetic data across scientific, industrial, and policy domains. While these developments open new possibilities for data analy…

Synthetic Data Generation

RealCause: Realistic Causal Inference Benchmarking

2020-11-30 · Brady Neal, Chin-wei Huang, Sunand Raghupathi

There are many different causal effect estimators in causal inference. However, it is unclear how to choose between these estimators because there is no ground-truth for causal effects. A commonly used option is to simul…

BenchmarkingCausal Inference