paper-with-me

Papers

WaKA: Data Attribution using K-Nearest Neighbors and Membership Privacy Principles

2024-11-02 · Patrick Mesana, Clément Bénesse, Hadrien Lautraite, Gilles Caporossi, Sébastien Gambs

In this paper, we introduce WaKA (Wasserstein K-nearest-neighbors Attribution), a novel attribution method that leverages principles from the LiRA (Likelihood Ratio Attack) framework and k-nearest neighbors classifiers (k-NN). WaKA efficiently measures the contribution of individual data points to the model's loss distribution, analyzing every possible k-NN that can be constructed using the training set, without requiring to sample subsets of the training set. WaKA is versatile and can be used a posteriori as a membership inference attack (MIA) to assess privacy risks or a priori for privacy influence measurement and data valuation. Thus, WaKA can be seen as bridging the gap between data attribution and membership inference attack (MIA) by providing a unified framework to distinguish between a data point's value and its privacy risk. For instance, we have shown that self-attribution values are more strongly correlated with the attack success rate than the contribution of a point to the model generalization. WaKA's different usage were also evaluated across diverse real-world datasets, demonstrating performance very close to LiRA when used as an MIA on k-NN classifiers, but with greater computational efficiency. Additionally, WaKA shows greater robustness than Shapley Values for data minimization tasks (removal or addition) on imbalanced datasets.

📄 PDF Abstract BibTeX arXiv:2411.01357

Code (0)

등록된 구현이 없습니다.

Tasks

Computational EfficiencyData ValuationInference AttackMembership Inference Attack

Methods 이 논문이 사용한 방법론

k-NN $k$-Nearest Neighbors is a clustering-based algorithm for classification and regression. It is a a type of instance-based learning as it does not attempt to construct a…

Similar Papers 제목 키워드 기반

Flexible K Nearest Neighbors Classifier: Derivation and Application for Ion-mobility Spectrometry-based Indoor Localization

2023-04-20 · Philipp Müller

The K Nearest Neighbors (KNN) classifier is widely used in many fields such as fingerprint-based localization or medicine. It determines the class membership of unlabelled sample based on the class memberships of the K l…

Indoor Localization

Label Noise Robustness for Domain-Agnostic Fair Corrections via Nearest Neighbors Label Spreading

2024-06-13 · Nathan Stromberg, Rohan Ayyagari, Sanmi Koyejo, Richard Nock 외

Last-layer retraining methods have emerged as an efficient framework for correcting existing base models. Within this framework, several methods have been proposed to deal with correcting models for subgroup fairness wit…

Fairness

Fuzzy k-Nearest Neighbors with monotonicity constraints: Moving towards the robustness of monotonic noise

2020-03-05 · Sergio González, Salvador García, Sheng-Tun Li, Robert John 외

This paper proposes a new model based on Fuzzy k-Nearest Neighbors for classification with monotonic constraints, Monotonic Fuzzy k-NN (MonFkNN). Real-life data-sets often do not comply with monotonic constraints due to …

Instance-based entropy fuzzy support vector machine for imbalanced data

2018-07-11 · Poongjin Cho, Minhyuk Lee, Woojin Chang

Imbalanced classification has been a major challenge for machine learning because many standard classifiers mainly focus on balanced datasets and tend to have biased results towards the majority class. We modify entropy …

BIG-bench Machine LearningDiversityimbalanced classification

Authorship attribution for Differences between Literary Texts by Bilingual Russian-French and Non-Bilingual French Authors

2023-03-01 · Margarita Makarova

Do bilingual Russian-French authors of the end of the twentieth century such as Andre\"i Makine, Val\'ery Afanassiev, Vladimir F\'edorovski, Iegor Gran, Luba Jurgenson have common stylistic traits in the novels they wrot…

Authorship Attribution