paper-with-me

홈 › Papers

Adversarial Neuron Pruning Purifies Backdoored Deep Models

2021-10-27 · NeurIPS 2021 12 · Dongxian Wu, Yisen Wang

As deep neural networks (DNNs) are growing larger, their requirements for computational resources become huge, which makes outsourcing training more popular. Training in a third-party platform, however, may introduce potential risks that a malicious trainer will return backdoored DNNs, which behave normally on clean samples but output targeted misclassifications whenever a trigger appears at the test time. Without any knowledge of the trigger, it is difficult to distinguish or recover benign DNNs from backdoored ones. In this paper, we first identify an unexpected sensitivity of backdoored DNNs, that is, they are much easier to collapse and tend to predict the target label on clean samples when their neurons are adversarially perturbed. Based on these observations, we propose a novel model repairing method, termed Adversarial Neuron Pruning (ANP), which prunes some sensitive neurons to purify the injected backdoor. Experiments show, even with only an extremely small amount of clean data (e.g., 1%), ANP effectively removes the injected backdoor without causing obvious performance degradation.

📄 PDF Abstract BibTeX arXiv:2110.14430

Code (2)

csdongxian/anp_backdoor 공식 구현 pytorch
csdongxian/csdongxian

Methods 이 논문이 사용한 방법론

Test 설명 없음
Pruning 설명 없음

Similar Papers 제목 키워드 기반

Defending against Backdoor Attack on Deep Neural Networks

2020-02-26 · Hao Cheng, Kaidi Xu, Sijia Liu, Pin-Yu Chen 외

Although deep neural networks (DNNs) have achieved a great success in various computer vision tasks, it is recently found that they are vulnerable to adversarial attacks. In this paper, we focus on the so-called \textit{…

Backdoor AttackData Poisoning

Test-Time Attention Purification for Backdoored Large Vision Language Models

2026-03-13 · Zhifang Zhang, Bojun Yang, Shuo He, Weitong Chen 외 arxiv

Despite the strong multimodal performance, large vision-language models (LVLMs) are vulnerable during fine-tuning to backdoor attacks, where adversaries insert trigger-embedded samples into the training data to implant b…

Fusing Pruned and Backdoored Models: Optimal Transport-based Data-free Backdoor Mitigation

2024-08-28 · Weilin Lin, Li Liu, Jianze Li, Hui Xiong

Backdoor attacks present a serious security threat to deep neuron networks (DNNs). Although numerous effective defense techniques have been proposed in recent years, they inevitably rely on the availability of either cle…

backdoor defense

Unveiling and Mitigating Backdoor Vulnerabilities based on Unlearning Weight Changes and Backdoor Activeness

2024-05-30 · Weilin Lin, Li Liu, Shaokui Wei, Jianze Li 외

The security threat of backdoor attacks is a central concern for deep neural networks (DNNs). Recently, without poisoned data, unlearning models with clean data and then learning a pruning mask have contributed to backdo…

backdoor defense

Reconstructive Neuron Pruning for Backdoor Defense

2023-05-24 · Yige Li, Xixiang Lyu, Xingjun Ma, Nodens Koren 외

Deep neural networks (DNNs) have been found to be vulnerable to backdoor attacks, raising security concerns about their deployment in mission-critical applications. While existing defense methods have demonstrated promis…

backdoor defense