paper-with-me

Papers

Removing the Trigger, Not the Backdoor: Alternative Triggers and Latent Backdoors

2026-03-10 · Gorka Abad, Ermes Franch, Stefanos Koffas, Stjepan Picek arxiv

Current backdoor defenses assume that neutralizing a known trigger removes the backdoor. We show this trigger-centric view is incomplete: \emph{alternative triggers}, patterns perceptually distinct from training triggers, reliably activate the same backdoor. We estimate the alternative trigger backdoor direction in feature space by contrasting clean and triggered representations, and then develop a feature-guided attack that jointly optimizes target prediction and directional alignment. First, we theoretically prove that alternative triggers exist and are an inevitable consequence of backdoor training. Then, we verify this empirically. Additionally, defenses that remove training triggers often leave backdoors intact, and alternative triggers can exploit the latent backdoor feature-space. Our findings motivate defenses targeting backdoor directions in representation space rather than input-space triggers.

📄 PDF Abstract BibTeX arXiv:2603.09772

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Forgetting to Forget: Attention Sink as A Gateway for Backdooring LLM Unlearning

2025-10-19 · Bingqi Shang, Yiwei Chen, Yihua Zhang, Bingquan Shen 외 arxiv

Large language model (LLM) unlearning is a key approach for removing undesired data, knowledge, or behaviors from pretrained models while retaining their general utility. Yet, with the rise of open-weight LLMs, we ask: c…

BadActs: A Universal Backdoor Defense in the Activation Space

2024-05-18 · Biao Yi, Sishuo Chen, Yiming Li, Tong Li 외

Backdoor attacks pose an increasingly severe security threat to Deep Neural Networks (DNNs) during their development stage. In response, backdoor sample purification has emerged as a promising defense mechanism, aiming t…

backdoor defense

BFClass: A Backdoor-free Text Classification Framework

2021-09-22 · Findings (EMNLP) 2021 11 · Zichao Li, Dheeraj Mekala, chengyu dong, Jingbo Shang

Backdoor attack introduces artificial vulnerabilities into the model by poisoning a subset of the training data via injecting triggers and modifying labels. Various trigger design strategies have been explored to attack …

Backdoor AttackClassificationLanguage ModelingLanguage Modelling+2

Is It Possible to Backdoor Face Forgery Detection with Natural Triggers?

2023-12-31 · Xiaoxuan Han, Songlin Yang, Wei Wang, Ziwen He 외

Deep neural networks have significantly improved the performance of face forgery detection models in discriminating Artificial Intelligent Generated Content (AIGC). However, their security is significantly threatened by …

Backdoor Attackbackdoor defense

DISTIL: Data-Free Inversion of Suspicious Trojan Inputs via Latent Diffusion

2025-07-30 · Hossein Mirzaei, Zeinab Taghavi, Sepehr Rezaee, Masoud Hadi 외 arxiv

Deep neural networks have demonstrated remarkable success across numerous tasks, yet they remain vulnerable to Trojan (backdoor) attacks, raising serious concerns about their safety in real-world mission-critical applica…

Object Detection