paper-with-me

홈 › Papers

Rethinking Backdoor Detection Evaluation for Language Models

2024-08-31 · Jun Yan, Wenjie Jacky Mo, Xiang Ren, Robin Jia

Backdoor attacks, in which a model behaves maliciously when given an attacker-specified trigger, pose a major security risk for practitioners who depend on publicly released language models. Backdoor detection methods aim to detect whether a released model contains a backdoor, so that practitioners can avoid such vulnerabilities. While existing backdoor detection methods have high accuracy in detecting backdoored models on standard benchmarks, it is unclear whether they can robustly identify backdoors in the wild. In this paper, we examine the robustness of backdoor detectors by manipulating different factors during backdoor planting. We find that the success of existing methods highly depends on how intensely the model is trained on poisoned data during backdoor planting. Specifically, backdoors planted with either more aggressive or more conservative training are significantly more difficult to detect than the default ones. Our results highlight a lack of robustness of existing backdoor detectors and the limitations in current benchmark construction.

📄 PDF Abstract BibTeX arXiv:2409.00399

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Rethinking Stealthiness of Backdoor Attack against NLP Models

2021-08-01 · ACL 2021 5 · Wenkai Yang, Yankai Lin, Peng Li, Jie zhou 외

Recent researches have shown that large natural language processing (NLP) models are vulnerable to a kind of security threat called the Backdoor Attack. Backdoor attacked models can achieve good performance on clean test…

Backdoor AttackData AugmentationSentiment AnalysisWord Embeddings

Rethinking Backdoor Attacks on Dataset Distillation: A Kernel Method Perspective

2023-11-28 · Ming-Yu Chung, Sheng-Yen Chou, Chia-Mu Yu, Pin-Yu Chen 외

Dataset distillation offers a potential means to enhance data efficiency in deep learning. Recent studies have shown its ability to counteract backdoor risks present in original training samples. In this study, we delve …

Backdoor AttackDataset Distillation

Rethinking Backdoor Attacks

2023-07-19 · Alaa Khaddaj, Guillaume Leclerc, Aleksandar Makelov, Kristian Georgiev 외

In a backdoor attack, an adversary inserts maliciously constructed backdoor examples into a training set to make the resulting model vulnerable to manipulation. Defending against such attacks typically involves viewing t…

Backdoor Attack

Rethinking the Backdoor Attacks' Triggers: A Frequency Perspective

2021-04-07 · ICCV 2021 10 · Yi Zeng, Won Park, Z. Morley Mao, Ruoxi Jia

Backdoor attacks have been considered a severe security threat to deep learning. Such attacks can make models perform abnormally on inputs with predefined triggers and still retain state-of-the-art performance on clean d…

Strong Backdoors for Default Logic

2016-02-19 · Johannes K. Fichte, Arne Meier, Irina Schindler

In this paper, we introduce a notion of backdoors to Reiter's propositional default logic and study structural properties of it. Also we consider the problems of backdoor detection (parameterised by the solution size) as…