paper-with-me

Papers

Backdooring Explainable Machine Learning

2022-04-20 · Maximilian Noppel, Lukas Peter, Christian Wressnegger

Explainable machine learning holds great potential for analyzing and understanding learning-based systems. These methods can, however, be manipulated to present unfaithful explanations, giving rise to powerful and stealthy adversaries. In this paper, we demonstrate blinding attacks that can fully disguise an ongoing attack against the machine learning model. Similar to neural backdoors, we modify the model's prediction upon trigger presence but simultaneously also fool the provided explanation. This enables an adversary to hide the presence of the trigger or point the explanation to entirely different portions of the input, throwing a red herring. We analyze different manifestations of such attacks for different explanation types in the image domain, before we resume to conduct a red-herring attack against malware classification.

📄 PDF Abstract BibTeX arXiv:2204.09498

Code (0)

등록된 구현이 없습니다.

Tasks

BIG-bench Machine LearningMalware Classification

Similar Papers 제목 키워드 기반

Dynamic Backdoor Attacks Against Machine Learning Models

2020-03-07 · Ahmed Salem, Rui Wen, Michael Backes, Shiqing Ma 외

Machine learning (ML) has made tremendous progress during the past decade and is being adopted in various critical real-world applications. However, recent research has shown that ML models are vulnerable to multiple sec…

Backdoor AttackBIG-bench Machine Learning

Dynamic Backdoor Attacks Against Deep Neural Networks

2021-01-01 · Ahmed Salem, Rui Wen, Michael Backes, Shiqing Ma 외

Current Deep Neural Network (DNN) backdooring attacks rely on adding static triggers (with fixed patterns and locations) on model inputs that are prone to detection. In this paper, we propose the first class of dynamic …

Semantic Shield: Defending Vision-Language Models Against Backdooring and Poisoning via Fine-grained Knowledge Alignment

2024-11-23 · CVPR 2024 1 · Alvi Md Ishmam, Christopher Thomas

In recent years there has been enormous interest in vision-language models trained using self-supervised objectives. However, the use of large-scale datasets scraped from the web for training also makes these models vuln…

Language ModelingLanguage Modelling

Forgetting to Forget: Attention Sink as A Gateway for Backdooring LLM Unlearning

2025-10-19 · Bingqi Shang, Yiwei Chen, Yihua Zhang, Bingquan Shen 외 arxiv

Large language model (LLM) unlearning is a key approach for removing undesired data, knowledge, or behaviors from pretrained models while retaining their general utility. Yet, with the rise of open-weight LLMs, we ask: c…

Curse or Redemption? How Data Heterogeneity Affects the Robustness of Federated Learning

2021-02-01 · Syed Zawad, Ahsan Ali, Pin-Yu Chen, Ali Anwar 외

Data heterogeneity has been identified as one of the key features in federated learning but often overlooked in the lens of robustness to adversarial attacks. This paper focuses on characterizing and understanding its im…

Federated Learning