paper-with-me

홈 › Papers

Comparing the Utility and Disclosure Risk of Synthetic Data with Samples of Microdata

2022-07-02 · Claire Little, Mark Elliot, Richard Allmendinger

Most statistical agencies release randomly selected samples of Census microdata, usually with sample fractions under 10% and with other forms of statistical disclosure control (SDC) applied. An alternative to SDC is data synthesis, which has been attracting growing interest, yet there is no clear consensus on how to measure the associated utility and disclosure risk of the data. The ability to produce synthetic Census microdata, where the utility and associated risks are clearly understood, could mean that more timely and wider-ranging access to microdata would be possible. This paper follows on from previous work by the authors which mapped synthetic Census data on a risk-utility (R-U) map. The paper presents a framework to measure the utility and disclosure risk of synthetic data by comparing it to samples of the original data of varying sample fractions, thereby identifying the sample fraction which has equivalent utility and risk to the synthetic data. Three commonly used data synthesis packages are compared with some interesting results. Further work is needed in several directions but the methodology looks very promising.

📄 PDF Abstract BibTeX arXiv:2207.03339

Code (1)

clairelittle/psd2022-comparing-utility-risk 공식 구현 pytorch

Similar Papers 제목 키워드 기반

Generating Synthetic Data with Locally Estimated Distributions for Disclosure Control

2022-10-03 · Ali Furkan Kalay

Sensitive datasets are often underutilized in research and industry due to privacy concerns, limiting the potential of valuable data-driven insights. Synthetic data generation presents a promising solution to address thi…

ClusteringHyperparameter OptimizationImputationModel Optimization+1

RAPID: Risk of Attribute Prediction-Induced Disclosure in Synthetic Microdata

2026-02-09 · Matthias Templ, Oscar Thees, Roman Müller arxiv

Statistical data anonymization increasingly relies on fully synthetic microdata, for which classical identity disclosure measures are less informative than an adversary's ability to infer sensitive attributes from releas…

Multi-objective evolutionary GAN for tabular data synthesis

2024-04-15 · Nian Ran, Bahrul Ilmi Nasution, Claire Little, Richard Allmendinger 외

Synthetic data has a key role to play in data sharing by statistical agencies and other generators of statistical data products. Generative Adversarial Networks (GANs), typically applied to image synthesis, are also a pr…

Image Generation

Understanding Latent Flow Models for Tabular Data Synthesis: Targets, Paths, and Sampling

2026-06-18 · Bahrul Ilmi Nasution arxiv

Synthetic tabular data enables microdata sharing in regulated domains, yet deploying continuous-time generative models requires balancing analytical utility, disclosure risk, and computational cost. Latent-space flow mod…

Protecting Vulnerable Voices: Synthetic Dataset Generation for Self-Disclosure Detection

2025-07-24 · Shalini Jangra, Suparna De, Nishanth Sastry, Saeed Fadaei arxiv

Social platforms such as Reddit have a network of communities of shared interests, with a prevalence of posts and comments from which one can infer users' Personal Information Identifiers (PIIs). While such self-disclosu…

Text GenerationText Detection