paper-with-me

Papers

SoK: Can Synthetic Images Replace Real Data? A Survey of Utility and Privacy of Synthetic Image Generation

2025-06-24 · Yunsung Chung, Yunbei Zhang, Nassir Marrouche, Jihun Hamm

Advances in generative models have transformed the field of synthetic image generation for privacy-preserving data synthesis (PPDS). However, the field lacks a comprehensive survey and comparison of synthetic image generation methods across diverse settings. In particular, when we generate synthetic images for the purpose of training a classifier, there is a pipeline of generation-sampling-classification which takes private training as input and outputs the final classifier of interest. In this survey, we systematically categorize existing image synthesis methods, privacy attacks, and mitigations along this generation-sampling-classification pipeline. To empirically compare diverse synthesis approaches, we provide a benchmark with representative generative methods and use model-agnostic membership inference attacks (MIAs) as a measure of privacy risk. Through this study, we seek to answer critical questions in PPDS: Can synthetic data effectively replace real data? Which release strategy balances utility and privacy? Do mitigations improve the utility-privacy tradeoff? Which generative models perform best across different scenarios? With a systematic evaluation of diverse methods, our study provides actionable insights into the utility-privacy tradeoffs of synthetic data generation methods and guides the decision on optimal data releasing strategies for real-world applications.

📄 PDF Abstract BibTeX arXiv:2506.19360

Code (0)

등록된 구현이 없습니다.

Tasks

Image GenerationPrivacy PreservingSurveySynthetic Data Generation

Similar Papers 제목 키워드 기반

Stochastic Parrots or Singing in Harmony? Testing Five Leading LLMs for their Ability to Replicate a Human Survey with Synthetic Data

2026-02-10 · Jason Miklian, Kristian Hoelscher, John E. Katsos arxiv

How well can AI-derived synthetic research data replicate the responses of human participants? An emerging literature has begun to engage with this question, which carries deep implications for organizational research pr…

Exploring the Potential of Synthetic Data to Replace Real Data

2024-08-26 · Hyungtae Lee, Yan Zhang, Heesung Kwon, Shuvra S. Bhattacharrya

The potential of synthetic data to replace real data creates a huge demand for synthetic data in data-hungry AI. This potential is even greater when synthetic data is used for training along with a small number of real i…

Using Zero-Shot LLM-Generated Survey Data for Geographically Explicit Population Synthesis

2026-04-23 · Taylor Anderson, Sara Von Hoene, Orhan Yagizer Cinar, Emma Von Hoene 외 arxiv

There is a growing interest in utilizing synthetic populations for a diverse range of applications. At the same time, we are witnessing a tremendous growth in artificial intelligence in all walks of life. This paper eval…

Railway Anomaly detection model using synthetic defect images generated by CycleGAN

2021-02-24 · Takuro Hoshi, Yohei Baba, Gaurang Gavai

Although training data is essential for machine learning, railway companies are facing difficulties in gathering adequate images of defective equipment due to their proactive replacement of would be defective equipment. …

Anomaly DetectionBIG-bench Machine LearningDefect Detection

MixDiff: Mixing Natural and Synthetic Images for Robust Self-Supervised Representations

2024-06-18 · Reza Akbarian Bafghi, Nidhin Harilal, Claire Monteleoni, Maziar Raissi

This paper introduces MixDiff, a new self-supervised learning (SSL) pre-training framework that combines real and synthetic images. Unlike traditional SSL methods that predominantly use real images, MixDiff uses a varian…

Image ClassificationSelf-Supervised Learning