paper-with-me

Papers

Understanding Impacts of Task Similarity on Backdoor Attack and Detection

2022-10-12 · Di Tang, Rui Zhu, XiaoFeng Wang, Haixu Tang, Yi Chen

With extensive studies on backdoor attack and detection, still fundamental questions are left unanswered regarding the limits in the adversary's capability to attack and the defender's capability to detect. We believe that answers to these questions can be found through an in-depth understanding of the relations between the primary task that a benign model is supposed to accomplish and the backdoor task that a backdoored model actually performs. For this purpose, we leverage similarity metrics in multi-task learning to formally define the backdoor distance (similarity) between the primary task and the backdoor task, and analyze existing stealthy backdoor attacks, revealing that most of them fail to effectively reduce the backdoor distance and even for those that do, still much room is left to further improve their stealthiness. So we further design a new method, called TSA attack, to automatically generate a backdoor model under a given distance constraint, and demonstrate that our new attack indeed outperforms existing attacks, making a step closer to understanding the attacker's limits. Most importantly, we provide both theoretic results and experimental evidence on various datasets for the positive correlation between the backdoor distance and backdoor detectability, demonstrating that indeed our task similarity analysis help us better understand backdoor risks and has the potential to identify more effective mitigations.

📄 PDF Abstract BibTeX arXiv:2210.06509

Code (0)

등록된 구현이 없습니다.

Tasks

Backdoor AttackMulti-Task Learning

Similar Papers 제목 키워드 기반

Backdoor Vectors: a Task Arithmetic View on Backdoor Attacks and Defenses

2025-10-09 · Stanisław Pawlak, Jan Dubiński, Daniel Marczak, Bartłomiej Twardowski arxiv

Model merging (MM) recently emerged as an effective method for combining large deep learning models. However, it poses significant security risks. Recent research shows that it is highly susceptible to backdoor attacks, …

A General Framework for Defending Against Backdoor Attacks via Influence Graph

2021-11-29 · Xiaofei Sun, Jiwei Li, Xiaoya Li, Ziyao Wang 외

In this work, we propose a new and general framework to defend against backdoor attacks, inspired by the fact that attack triggers usually follow a \textsc{specific} type of attacking pattern, and therefore, poisoned tra…

Multi-target Backdoor Attacks for Code Pre-trained Models

2023-06-14 · Yanzhou Li, Shangqing Liu, Kangjie Chen, Xiaofei Xie 외

Backdoor attacks for neural code models have gained considerable attention due to the advancement of code intelligence. However, most existing works insert triggers into task-specific data for code-related downstream tas…

Code GenerationRepresentation Learning

Uncovering and Aligning Anomalous Attention Heads to Defend Against NLP Backdoor Attacks

2025-11-16 · Haotian Jin, Yang Li, Haihui Fan, Lin Shen 외 arxiv

Backdoor attacks pose a serious threat to the security of large language models (LLMs), causing them to exhibit anomalous behavior under specific trigger conditions. The design of backdoor triggers has evolved from fixed…

Unlearning Backdoor Threats: Enhancing Backdoor Defense in Multimodal Contrastive Learning via Local Token Unlearning

2024-03-24 · Siyuan Liang, Kuanrong Liu, Jiajun Gong, Jiawei Liang 외

Multimodal contrastive learning has emerged as a powerful paradigm for building high-quality features using the complementary strengths of various data modalities. However, the open nature of such systems inadvertently i…

backdoor defenseContrastive Learning