paper-with-me

홈 › Papers

Generating Higher-Fidelity Synthetic Datasets with Privacy Guarantees

2020-03-02 · Aleksei Triastcyn, Boi Faltings

This paper considers the problem of enhancing user privacy in common machine learning development tasks, such as data annotation and inspection, by substituting the real data with samples form a generative adversarial network. We propose employing Bayesian differential privacy as the means to achieve a rigorous theoretical guarantee while providing a better privacy-utility trade-off. We demonstrate experimentally that our approach produces higher-fidelity samples, compared to prior work, allowing to (1) detect more subtle data errors and biases, and (2) reduce the need for real data labelling by achieving high accuracy when training directly on artificial samples.

📄 PDF Abstract BibTeX arXiv:2003.00997

Code (0)

등록된 구현이 없습니다.

Tasks

BIG-bench Machine LearningGenerative Adversarial Network

Similar Papers 제목 키워드 기반

Defining 'Good': Evaluation Framework for Synthetic Smart Meter Data

2024-07-16 · Sheng Chai, Gus Chadney, Charlot Avery, Phil Grunewald 외

Access to granular demand data is essential for the net zero transition; it allows for accurate profiling and active demand management as our reliance on variable renewable generation increases. However, public release o…

Generating tabular datasets under differential privacy

2023-08-28 · Gianluca Truda

Machine Learning (ML) is accelerating progress across fields and industries, but relies on accessible and high-quality training data. Some of the most important datasets are found in biomedical and financial domains in t…

Synthetic Data Generation

Synthetic Data Generation and Differential Privacy using Tensor Networks' Matrix Product States (MPS)

2025-08-08 · Alejandro Moreno R., Desale Fentaw, Samuel Palmer, Raúl Salles de Padua 외 arxiv

Synthetic data generation is a key technique in modern artificial intelligence, addressing data scarcity, privacy constraints, and the need for diverse datasets in training robust models. In this work, we propose a metho…

Synthetic Data Generation

No Free Lunch for Synthetic Images under Data Scarcity Conditions

2026-06-01 · Borja Arroyo Galende, Alejandro Almodóvar, Patricia A. Apellániz, Juan Parras 외 arxiv

This study investigates the trade-offs between fidelity, privacy, and utility in synthetic data generation under conditions of data scarcity and privacy sensitivity. We propose an evaluation framework that jointly assess…

Synthetic Data Generation

Generative Correlation Manifolds: Generating Synthetic Data with Preserved Higher-Order Correlations

2025-10-24 · Jens E. d'Hondt, Wieger R. Punter, Odysseas Papapetrou arxiv

The increasing need for data privacy and the demand for robust machine learning models have fueled the development of synthetic data generation techniques. However, current methods often succeed in replicating simple sum…

Synthetic Data Generation