paper-with-me

홈 › Papers

Harnessing Synthetic Data from Generative AI for Statistical Inference

2026-03-05 · Ahmad Abdel-Azim, Ruoyu Wang, Xihong Lin arxiv

The emergence of generative AI models has dramatically expanded the availability and use of synthetic data across scientific, industrial, and policy domains. While these developments open new possibilities for data analysis, they also raise fundamental statistical questions about when synthetic data can be used in a valid, reliable, and principled manner. This paper reviews the current landscape of synthetic data generation and use from a statistical perspective, with the goal of clarifying the assumptions under which synthetic data can meaningfully support downstream discovery, inference, and prediction. We survey major classes of modern generative models, their intended use cases, and the benefits they offer, while also highlighting their limitations and characteristic failure modes. We additionally examine common pitfalls that arise when synthetic data are treated as surrogates for real observations, including biases from model misspecification, attenuated uncertainty, and difficulties in generalization. Building on these insights, we discuss emerging frameworks for the principled use of synthetic data. We conclude with practical recommendations, open problems, and cautions intended to guide both method developers and applied researchers.

📄 PDF Abstract BibTeX arXiv:2603.05396

Code (0)

등록된 구현이 없습니다.

Tasks

Synthetic Data Generation

Similar Papers 제목 키워드 기반

Boosting Data Analytics With Synthetic Volume Expansion

2023-10-27 · Xiaotong Shen, Yifei Liu, Rex Shen

Synthetic data generation, a cornerstone of Generative Artificial Intelligence, promotes a paradigm shift in data science by addressing data scarcity and privacy while enabling unprecedented performance. As synthetic dat…

Sentiment AnalysisSynthetic Data GenerationTransfer Learning

The Real Deal Behind the Artificial Appeal: Inferential Utility of Tabular Synthetic Data

2023-12-13 · Alexander Decruyenaere, Heidelinde Dehaene, Paloma Rabaey, Christiaan Polet 외

Recent advances in generative models facilitate the creation of synthetic data to be made available for research in privacy-sensitive contexts. However, the analysis of synthetic data raises a unique set of methodologica…

Mitigating Statistical Bias within Differentially Private Synthetic Data

2021-08-24 · Sahra Ghalebikesabi, Harrison Wilde, Jack Jewson, Arnaud Doucet 외

Increasing interest in privacy-preserving machine learning has led to new and evolved approaches for generating private synthetic data from undisclosed real data. However, mechanisms of privacy preservation can significa…

Privacy Preserving

A Linear Reconstruction Approach for Attribute Inference Attacks against Synthetic Data

2023-01-24 · Meenatchi Sundaram Muthu Selva Annamalai, Andrea Gadotti, Luc Rocher

Recent advances in synthetic data generation (SDG) have been hailed as a solution to the difficult problem of sharing sensitive data while protecting privacy. SDG aims to learn statistical properties of real data in orde…

AttributeInference AttackSynthetic Data Generation

Generative Models for Simulating Mobility Trajectories

2018-11-30 · Vaibhav Kulkarni, Natasa Tagasovska, Thibault Vatter, Benoit Garbinato

Mobility datasets are fundamental for evaluating algorithms pertaining to geographic information systems and facilitating experimental reproducibility. But privacy implications restrict sharing such datasets, as even agg…

Semantic SimilaritySemantic Textual Similarity