paper-with-me

홈 › Papers

Privacy for Free: How does Dataset Condensation Help Privacy?

2022-06-01 · Tian Dong, Bo Zhao, Lingjuan Lyu

To prevent unintentional data leakage, research community has resorted to data generators that can produce differentially private data for model training. However, for the sake of the data privacy, existing solutions suffer from either expensive training cost or poor generalization performance. Therefore, we raise the question whether training efficiency and privacy can be achieved simultaneously. In this work, we for the first time identify that dataset condensation (DC) which is originally designed for improving training efficiency is also a better solution to replace the traditional data generators for private data generation, thus providing privacy for free. To demonstrate the privacy benefit of DC, we build a connection between DC and differential privacy, and theoretically prove on linear feature extractors (and then extended to non-linear feature extractors) that the existence of one sample has limited impact ($O(m/n)$) on the parameter distribution of networks trained on $m$ samples synthesized from $n (n \gg m)$ raw samples by DC. We also empirically validate the visual privacy and membership privacy of DC-synthesized data by launching both the loss-based and the state-of-the-art likelihood-based membership inference attacks. We envision this work as a milestone for data-efficient and privacy-preserving machine learning.

📄 PDF Abstract BibTeX arXiv:2206.00240

Code (1)

Guang000/Awesome-Dataset-Distillation

Tasks

Dataset CondensationPrivacy Preserving

Similar Papers 제목 키워드 기반

No Free Lunch in "Privacy for Free: How does Dataset Condensation Help Privacy"

2022-09-29 · Nicholas Carlini, Vitaly Feldman, Milad Nasr

New methods designed to preserve data privacy require careful scrutiny. Failure to preserve privacy is hard to detect, and yet can lead to catastrophic results when a system implementing a ``privacy-preserving'' method i…

Dataset CondensationPrivacy Preserving

Connect the dots: Dataset Condensation, Differential Privacy, and Adversarial Uncertainty

2024-02-16 · Kenneth Odoh

Our work focuses on understanding the underpinning mechanism of dataset condensation by drawing connections with ($\epsilon$, $\delta$)-differential privacy where the optimal noise, $\epsilon$, is chosen by adversarial u…

Dataset CondensationNoise Estimation

DCFL: Non-IID awareness Data Condensation aided Federated Learning

2023-12-21 · Shaohan Sha, YaFeng Sun

Federated learning is a decentralized learning paradigm wherein a central server trains a global model iteratively by utilizing clients who possess a certain amount of private datasets. The challenge lies in the fact tha…

Dataset CondensationFederated Learning

Improving Clinical Dataset Condensation with Mode Connectivity-based Trajectory Surrogates

2025-10-07 · Pafue Christy Nganjimi, Andrew Soltan, Danielle Belgrave, Lei Clifton 외 arxiv

Dataset condensation (DC) enables the creation of compact, privacy-preserving synthetic datasets that can match the utility of real patient records, supporting democratised access to highly regulated clinical data for de…

Is Less More? Exploring Token Condensation as Training-free Adaptation for CLIP

2024-10-16 · Zixin Wang, Dong Gong, Sen Wang, Zi Huang 외

Contrastive language-image pre-training (CLIP) has shown remarkable generalization ability in image classification. However, CLIP sometimes encounters performance drops on downstream datasets during zero-shot inference. …

image-classificationImage ClassificationTest-time Adaptation