paper-with-me

홈 › Papers

How to Sift Out a Clean Data Subset in the Presence of Data Poisoning?

2022-10-12 · Yi Zeng, Minzhou Pan, Himanshu Jahagirdar, Ming Jin, Lingjuan Lyu, Ruoxi Jia

Given the volume of data needed to train modern machine learning models, external suppliers are increasingly used. However, incorporating external data poses data poisoning risks, wherein attackers manipulate their data to degrade model utility or integrity. Most poisoning defenses presume access to a set of clean data (or base set). While this assumption has been taken for granted, given the fast-growing research on stealthy poisoning attacks, a question arises: can defenders really identify a clean subset within a contaminated dataset to support defenses? This paper starts by examining the impact of poisoned samples on defenses when they are mistakenly mixed into the base set. We analyze five defenses and find that their performance deteriorates dramatically with less than 1% poisoned points in the base set. These findings suggest that sifting out a base set with high precision is key to these defenses' performance. Motivated by these observations, we study how precise existing automated tools and human inspection are at identifying clean data in the presence of data poisoning. Unfortunately, neither effort achieves the precision needed. Worse yet, many of the outcomes are worse than random selection. In addition to uncovering the challenge, we propose a practical countermeasure, Meta-Sift. Our method is based on the insight that existing attacks' poisoned samples shifts from clean data distributions. Hence, training on the clean portion of a dataset and testing on the corrupted portion will result in high prediction loss. Leveraging the insight, we formulate a bilevel optimization to identify clean data and further introduce a suite of techniques to improve efficiency and precision. Our evaluation shows that Meta-Sift can sift a clean base set with 100% precision under a wide range of poisoning attacks. The selected base set is large enough to give rise to successful defenses.

📄 PDF Abstract BibTeX arXiv:2210.06516

Code (1)

ruoxi-jia-group/meta-sift 공식 구현 pytorch

Tasks

Bilevel OptimizationData Poisoning

Methods 이 논문이 사용한 방법론

BASE 설명 없음

Similar Papers 제목 키워드 기반

Robust Palm-Vein Recognition Using the MMD Filter: Improving SIFT-Based Feature Matching

2025-03-03 · Kaveen Perera, Fouad Khelifi, Ammar Belatreche

A major challenge with palm vein images is that slight movements of the fingers and thumb, or variations in hand posture, can stretch the skin in different areas and alter the vein patterns. This can result in an infinit…

Exact Unlearning of Finetuning Data via Model Merging at Scale

2025-04-06 · Kevin Kuo, Amrith Setlur, Kartik Srinivas, aditi raghunathan 외

Approximate unlearning has gained popularity as an approach to efficiently update an LLM so that it behaves (roughly) as if it was not trained on a subset of data to begin with. However, existing methods are brittle in p…

ASM: Adaptive Sample Mining for In-The-Wild Facial Expression Recognition

2023-10-09 · Ziyang Zhang, Xiao Sun, Liuwei An, Meng Wang

Given the similarity between facial expression categories, the presence of compound facial expressions, and the subjectivity of annotators, facial expression recognition (FER) datasets often suffer from ambiguity and noi…

Facial Expression RecognitionFacial Expression Recognition (FER)

Not All Features Are Equal: Discovering Essential Features for Preserving Prediction Privacy

2020-03-26 · Fatemehsadat Mireshghallah, Mohammadkazem Taram, Ali Jalali, Ahmed Taha Elthakeb 외

When receiving machine learning services from the cloud, the provider does not need to receive all features; in fact, only a subset of the features are necessary for the target prediction task. Discerning this subset is …

All

An analysis of the factors affecting keypoint stability in scale-space

2015-11-26 · Ives Rey-Otero, Jean-Michel Morel, Mauricio Delbracio

The most popular image matching algorithm SIFT, introduced by D. Lowe a decade ago, has proven to be sufficiently scale invariant to be used in numerous applications. In practice, however, scale invariance may be weakene…

Keypoint Detection