paper-with-me

홈 › Papers

Improving the Generation and Evaluation of Synthetic Data for Downstream Medical Causal Inference

2025-10-21 · Harry Amad, Zhaozhi Qian, Dennis Frauen, Julianna Piskorz, Stefan Feuerriegel, Mihaela van der Schaar arxiv

Causal inference is essential for developing and evaluating medical interventions, yet real-world medical datasets are often difficult to access due to regulatory barriers. This makes synthetic data a potentially valuable asset that enables these medical analyses, along with the development of new inference methods themselves. Generative models can produce synthetic data that closely approximate real data distributions, yet existing methods do not consider the unique challenges that downstream causal inference tasks, and specifically those focused on treatments, pose. We establish a set of desiderata that synthetic data containing treatments should satisfy to maximise downstream utility: preservation of (i) the covariate distribution, (ii) the treatment assignment mechanism, and (iii) the outcome generation mechanism. Based on these desiderata, we propose a set of evaluation metrics to assess such synthetic data. Finally, we present STEAM: a novel method for generating Synthetic data for Treatment Effect Analysis in Medicine that mimics the data-generating process of data containing treatments and optimises for our desiderata. We empirically demonstrate that STEAM achieves state-of-the-art performance across our metrics as compared to existing generative models, particularly as the complexity of the true data-generating process increases.

📄 PDF Abstract BibTeX arXiv:2510.18768

Code (0)

등록된 구현이 없습니다.

Tasks

Causal Inference

Similar Papers 제목 키워드 기반

Generation of Synthetic Clinical Text: A Systematic Review

2025-07-24 · Basel Alshaikhdeeb, Ahmed Abdelmonem Hemedan, Soumyabrata Ghosh, Irina Balaur 외 arxiv

Generating clinical synthetic text represents an effective solution for common clinical NLP issues like sparsity and privacy. This paper aims to conduct a systematic review on generating synthetic medical free-text by fo…

Generating Synthetic Free-text Medical Records with Low Re-identification Risk using Masked Language Modeling

2024-09-15 · Samuel Belkadi, Libo Ren, Nicolo Micheletti, Lifeng Han 외

The vast amount of available medical records has the potential to improve healthcare and biomedical research. However, privacy restrictions make these data accessible for internal use only. Recent works have addressed th…

Causal Language ModelingDe-identificationDiversityLanguage Modeling+4

Beyond a Single Mode: GAN Ensembles for Diverse Medical Data Generation

2025-03-31 · Lorenzo Tronchin, Tommy Löfstedt, Paolo Soda, Valerio Guarrasi

The advancement of generative AI, particularly in medical imaging, confronts the trilemma of ensuring high fidelity, diversity, and efficiency in synthetic data generation. While Generative Adversarial Networks (GANs) ha…

DiagnosticDiversitySynthetic Data Generation

Patient-Zero: Scaling Synthetic Patient Agents to Real-World Distributions without Real Patient Data

2025-09-14 · Yunghwei Lai, Ziyue Wang, Weizhi Ma, Yang Liu arxiv

Synthetic data generation with Large Language Models (LLMs) has emerged as a promising solution in the medical domain to mitigate data scarcity and privacy constraints. However, existing approaches remain constrained by …

Synthetic Data Generation

Vision-Language Synthetic Data Enhances Echocardiography Downstream Tasks

2024-03-28 · Pooria Ashrafian, Milad Yazdani, Moein Heidari, Dena Shahriari 외

High-quality, large-scale data is essential for robust deep learning models in medical applications, particularly ultrasound image analysis. Diffusion models facilitate high-fidelity medical image generation, reducing th…

Image GenerationMedical Image Generation