paper-with-me

홈 › Papers

REFINE: Inversion-Free Backdoor Defense via Model Reprogramming

2025-02-22 · Yukun Chen, Shuo Shao, Enhao Huang, Yiming Li, Pin-Yu Chen, Zhan Qin, Kui Ren

Backdoor attacks on deep neural networks (DNNs) have emerged as a significant security threat, allowing adversaries to implant hidden malicious behaviors during the model training phase. Pre-processing-based defense, which is one of the most important defense paradigms, typically focuses on input transformations or backdoor trigger inversion (BTI) to deactivate or eliminate embedded backdoor triggers during the inference process. However, these methods suffer from inherent limitations: transformation-based defenses often fail to balance model utility and defense performance, while BTI-based defenses struggle to accurately reconstruct trigger patterns without prior knowledge. In this paper, we propose REFINE, an inversion-free backdoor defense method based on model reprogramming. REFINE consists of two key components: \textbf{(1)} an input transformation module that disrupts both benign and backdoor patterns, generating new benign features; and \textbf{(2)} an output remapping module that redefines the model's output domain to guide the input transformations effectively. By further integrating supervised contrastive loss, REFINE enhances the defense capabilities while maintaining model utility. Extensive experiments on various benchmark datasets demonstrate the effectiveness of our REFINE and its resistance to potential adaptive attacks.

📄 PDF Abstract BibTeX arXiv:2502.18508

Code (2)

thuyimingli/backdoorbox 공식 구현 pytorch
whitolfchen/refine 공식 구현 pytorch

Tasks

backdoor defense

Similar Papers 제목 키워드 기반

Turning a Curse into a Blessing: Enabling In-Distribution-Data-Free Backdoor Removal via Stabilized Model Inversion

2022-06-14 · Si Chen, Yi Zeng, Jiachen T. Wang, Won Park 외

Many backdoor removal techniques in machine learning models require clean in-distribution data, which may not always be available due to proprietary datasets. Model inversion techniques, often considered privacy threats,…

BAN: Detecting Backdoors Activated by Adversarial Neuron Noise

2024-05-30 · Xiaoyun Xu, Zhuoran Liu, Stefanos Koffas, Shujian Yu 외

Backdoor attacks on deep learning represent a recent threat that has gained significant attention in the research community. Backdoor defenses are mainly based on backdoor inversion, which has been shown to be generic, m…

CEPA: Consensus Embedded Perturbation for Agnostic Detection and Inversion of Backdoors

2024-02-03 · Guangmingmei Yang, Xi Li, Hang Wang, David J. Miller 외

A variety of defenses have been proposed against Trojans planted in (backdoor attacks on) deep neural network (DNN) classifiers. Backdoor-agnostic methods seek to reliably detect and/or to mitigate backdoors irrespective…

image-classificationImage Classification

A Dual-Purpose Framework for Backdoor Defense and Backdoor Amplification in Diffusion Models

2025-02-26 · Vu Tuan Truong, Long Bao Le

Diffusion models have emerged as state-of-the-art generative frameworks, excelling in producing high-quality multi-modal samples. However, recent studies have revealed their vulnerability to backdoor attacks, where backd…

Backdoor Attackbackdoor defenseDenoising

DISTIL: Data-Free Inversion of Suspicious Trojan Inputs via Latent Diffusion

2025-07-30 · Hossein Mirzaei, Zeinab Taghavi, Sepehr Rezaee, Masoud Hadi 외 arxiv

Deep neural networks have demonstrated remarkable success across numerous tasks, yet they remain vulnerable to Trojan (backdoor) attacks, raising serious concerns about their safety in real-world mission-critical applica…

Object Detection