paper-with-me

홈 › Papers

Adversarial attacks and defenses in explainable artificial intelligence: A survey

2023-06-06 · Hubert Baniecki, Przemyslaw Biecek

Explainable artificial intelligence (XAI) methods are portrayed as a remedy for debugging and trusting statistical and deep learning models, as well as interpreting their predictions. However, recent advances in adversarial machine learning (AdvML) highlight the limitations and vulnerabilities of state-of-the-art explanation methods, putting their security and trustworthiness into question. The possibility of manipulating, fooling or fairwashing evidence of the model's reasoning has detrimental consequences when applied in high-stakes decision-making and knowledge discovery. This survey provides a comprehensive overview of research concerning adversarial attacks on explanations of machine learning models, as well as fairness metrics. We introduce a unified notation and taxonomy of methods facilitating a common ground for researchers and practitioners from the intersecting research fields of AdvML and XAI. We discuss how to defend against attacks and design robust interpretation methods. We contribute a list of existing insecurities in XAI and outline the emerging research directions in adversarial XAI (AdvXAI). Future work should address improving explanation methods and evaluation protocols to take into account the reported safety issues.

📄 PDF Abstract BibTeX arXiv:2306.06123

Code (1)

hbaniecki/adversarial-explainable-ai 공식 구현 pytorch

Tasks

Decision MakingExplainable artificial intelligenceExplainable Artificial Intelligence (XAI)FairnessSurvey

Similar Papers 제목 키워드 기반

Robust Intrusion Detection System with Explainable Artificial Intelligence

2025-03-07 · Betül Güvenç Paltun, Ramin Fuladi, Rim El Malki

Machine learning (ML) models serve as powerful tools for threat detection and mitigation; however, they also introduce potential new risks. Adversarial input can exploit these models through standard interfaces, thus cre…

Explainable artificial intelligenceExplainable Artificial Intelligence (XAI)Intrusion Detection

RAB$^2$-DEF: Dynamic and explainable defense against adversarial attacks in Federated Learning to fair poor clients

2024-10-10 · Nuria Rodríguez-Barroso, M. Victoria Luzón, Francisco Herrera

At the same time that artificial intelligence is becoming popular, concern and the need for regulation is growing, including among other requirements the data privacy. In this context, Federated Learning is proposed as a…

FairnessFederated Learning

Advancing the Research and Development of Assured Artificial Intelligence and Machine Learning Capabilities

2020-09-24 · Tyler J. Shipp, Daniel J. Clouse, Michael J. De Lucia, Metin B. Ahiskali 외

Artificial intelligence (AI) and machine learning (ML) have become increasingly vital in the development of novel defense and intelligence capabilities across all domains of warfare. An adversarial AI (A2I) and adversari…

BIG-bench Machine Learning

Enhancing Adversarial Robustness via Score-Based Optimization

2023-07-10 · NeurIPS 2023 11 · Boya Zhang, Weijian Luo, Zhihua Zhang

Adversarial attacks have the potential to mislead deep neural network classifiers by introducing slight perturbations. Developing algorithms that can mitigate the effects of these attacks is crucial for ensuring the safe…

Adversarial DefenseAdversarial Robustness

Enhancing Adversarial Robustness of IoT Intrusion Detection via SHAP-Based Attribution Fingerprinting

2025-11-09 · Dilli Prasad Sharma, Liang Xue, Xiaowei Sun, Xiaodong Lin 외 arxiv

The rapid proliferation of Internet of Things (IoT) devices has transformed numerous industries by enabling seamless connectivity and data-driven automation. However, this expansion has also exposed IoT networks to incre…

Adversarial RobustnessIntrusion Detection