Sniper GMMs: Structured Gaussian mixtures poison ML on large n small p data with high efficacy
We propose a method for structured learning of Gaussian mixtures with low KL-divergence from target mixture models that in turn model the raw data. We show that samples from these structured distributions are highly effective and evasive in poisoning training datasets of popular machine learning training pipelines such as neural networks, XGBoost and random forests. Such attacks are especially destructive given the current uptrends towards distributed machine learning with several untrusted client devices that provide their data to servers and cloud service providers for privacy preserving distributed machine learning. In current day and age of machine learning, Gaussian mixtures are perceived to be an older/classical technique in practice, although they are still actively studied from a theoretical perspective. Therefore it is quite interesting to see that they can be highly effective in performing data poisoning attacks on complex ML pipelines if learned with the right structural constraints.
Code (0)
등록된 구현이 없습니다.
Tasks
BIG-bench Machine LearningData PoisoningPrivacy PreservingSimilar Papers 제목 키워드 기반
Mixtures of Gaussians are Privately Learnable with a Polynomial Number of Samples
We study the problem of estimating mixtures of Gaussians under the constraint of differential privacy (DP). Our main result is that $\text{poly}(k,d,1/\alpha,1/\varepsilon,\log(1/\delta))$ samples are sufficient to estim…
Projection pursuit based on Gaussian mixtures and evolutionary algorithms
We propose a projection pursuit (PP) algorithm based on Gaussian mixture models (GMMs). The negentropy obtained from a multivariate density estimated by GMMs is adopted as the PP index to be maximised. For a fixed dimens…
Density EstimationEvolutionary AlgorithmsEvaluating generative networks using Gaussian mixtures of image features
We develop a measure for evaluating the performance of generative networks given two sets of images. A popular performance measure currently used to do this is the Fr\'echet Inception Distance (FID). FID assumes that ima…
Sublinear Variational Optimization of Gaussian Mixture Models with Millions to Billions of Parameters
Gaussian Mixture Models (GMMs) range among the most frequently used machine learning models. However, training large, general GMMs becomes computationally prohibitive for datasets with many data points $N$ of high-dimens…
CPUAgnostic Private Density Estimation for GMMs via List Global Stability
We consider the problem of private density estimation for mixtures of unrestricted high dimensional Gaussians in the agnostic setting. We prove the first upper bound on the sample complexity of this problem. Previously, …
Density Estimation