Generating Higher-Fidelity Synthetic Datasets with Privacy Guarantees
This paper considers the problem of enhancing user privacy in common machine learning development tasks, such as data annotation and inspection, by substituting the real data with samples form a generative adversarial network. We propose employing Bayesian differential privacy as the means to achieve a rigorous theoretical guarantee while providing a better privacy-utility trade-off. We demonstrate experimentally that our approach produces higher-fidelity samples, compared to prior work, allowing to (1) detect more subtle data errors and biases, and (2) reduce the need for real data labelling by achieving high accuracy when training directly on artificial samples.
Code (0)
등록된 구현이 없습니다.
Tasks
BIG-bench Machine LearningGenerative Adversarial NetworkSimilar Papers 제목 키워드 기반
Defining 'Good': Evaluation Framework for Synthetic Smart Meter Data
Access to granular demand data is essential for the net zero transition; it allows for accurate profiling and active demand management as our reliance on variable renewable generation increases. However, public release o…
Generating tabular datasets under differential privacy
Machine Learning (ML) is accelerating progress across fields and industries, but relies on accessible and high-quality training data. Some of the most important datasets are found in biomedical and financial domains in t…
Synthetic Data GenerationSynthetic Data Generation and Differential Privacy using Tensor Networks' Matrix Product States (MPS)
Synthetic data generation is a key technique in modern artificial intelligence, addressing data scarcity, privacy constraints, and the need for diverse datasets in training robust models. In this work, we propose a metho…
Synthetic Data GenerationNo Free Lunch for Synthetic Images under Data Scarcity Conditions
This study investigates the trade-offs between fidelity, privacy, and utility in synthetic data generation under conditions of data scarcity and privacy sensitivity. We propose an evaluation framework that jointly assess…
Synthetic Data GenerationGenerative Correlation Manifolds: Generating Synthetic Data with Preserved Higher-Order Correlations
The increasing need for data privacy and the demand for robust machine learning models have fueled the development of synthetic data generation techniques. However, current methods often succeed in replicating simple sum…
Synthetic Data Generation