paper-with-me

홈 › Papers

Robin Hood and Matthew Effects: Differential Privacy Has Disparate Impact on Synthetic Data

2021-09-23 · Georgi Ganev, Bristena Oprisanu, Emiliano De Cristofaro

Generative models trained with Differential Privacy (DP) can be used to generate synthetic data while minimizing privacy risks. We analyze the impact of DP on these models vis-a-vis underrepresented classes/subgroups of data, specifically, studying: 1) the size of classes/subgroups in the synthetic data, and 2) the accuracy of classification tasks run on them. We also evaluate the effect of various levels of imbalance and privacy budgets. Our analysis uses three state-of-the-art DP models (PrivBayes, DP-WGAN, and PATE-GAN) and shows that DP yields opposite size distributions in the generated synthetic data. It affects the gap between the majority and minority classes/subgroups; in some cases by reducing it (a "Robin Hood" effect) and, in others, by increasing it (a "Matthew" effect). Either way, this leads to (similar) disparate impacts on the accuracy of classification tasks on the synthetic data, affecting disproportionately more the underrepresented subparts of the data. Consequently, when training models on synthetic data, one might incur the risk of treating different subpopulations unevenly, leading to unreliable or unfair conclusions.

📄 PDF Abstract BibTeX arXiv:2109.11429

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

ReuseKNN: Neighborhood Reuse for Differentially-Private KNN-Based Recommendations

2022-06-23 · Peter Müllner, Elisabeth Lex, Markus Schedl, Dominik Kowald

User-based KNN recommender systems (UserKNN) utilize the rating data of a target user's k nearest neighbors in the recommendation process. This, however, increases the privacy risk of the neighbors since their rating dat…

Recommendation Systems

A Neighbourhood-Aware Differential Privacy Mechanism for Static Word Embeddings

2023-09-19 · Danushka Bollegala, Shuichi Otake, Tomoya Machide, Ken-ichi Kawarabayashi

We propose a Neighbourhood-Aware Differential Privacy (NADP) mechanism considering the neighbourhood of a word in a pretrained static word embedding space to determine the minimal amount of noise required to guarantee a …

Word Embeddings

Eternal Sunshine of the Spotless Net: Selective Forgetting in Deep Networks

2019-11-12 · CVPR 2020 6 · Aditya Golatkar, Alessandro Achille, Stefano Soatto

We explore the problem of selectively forgetting a particular subset of the data used for training a deep neural network. While the effects of the data to be forgotten can be hidden from the output of the network, insigh…

Differentially Private Estimation of Heterogeneous Causal Effects

2022-02-22 · Fengshi Niu, Harsha Nori, Brian Quistorff, Rich Caruana 외

Estimating heterogeneous treatment effects in domains such as healthcare or social science often involves sensitive data where protecting privacy is important. We introduce a general meta-algorithm for estimating conditi…

On the Convergence of Differentially-Private Fine-tuning: To Linearly Probe or to Fully Fine-tune?

2024-02-29 · Shuqi Ke, Charlie Hou, Giulia Fanti, Sewoong Oh

Differentially private (DP) machine learning pipelines typically involve a two-phase process: non-private pre-training on a public dataset, followed by fine-tuning on private data using DP optimization techniques. In the…