paper-with-me

홈 › Papers

Do Gradient-based Explanations Tell Anything About Adversarial Robustness to Android Malware?

2020-05-04 · Marco Melis, Michele Scalas, Ambra Demontis, Davide Maiorca, Battista Biggio, Giorgio Giacinto, Fabio Roli

While machine-learning algorithms have demonstrated a strong ability in detecting Android malware, they can be evaded by sparse evasion attacks crafted by injecting a small set of fake components, e.g., permissions and system calls, without compromising intrusive functionality. Previous work has shown that, to improve robustness against such attacks, learning algorithms should avoid overemphasizing few discriminant features, providing instead decisions that rely upon a large subset of components. In this work, we investigate whether gradient-based attribution methods, used to explain classifiers' decisions by identifying the most relevant features, can be used to help identify and select more robust algorithms. To this end, we propose to exploit two different metrics that represent the evenness of explanations, and a new compact security measure called Adversarial Robustness Metric. Our experiments conducted on two different datasets and five classification algorithms for Android malware detection show that a strong connection exists between the uniformity of explanations and adversarial robustness. In particular, we found that popular techniques like Gradient*Input and Integrated Gradients are strongly correlated to security when applied to both linear and nonlinear detectors, while more elementary explanation techniques like the simple Gradient do not provide reliable information about the robustness of such classifiers.

📄 PDF Abstract BibTeX arXiv:2005.01452

Code (0)

등록된 구현이 없습니다.

Tasks

Adversarial RobustnessAndroid Malware DetectionMalware Detection

Similar Papers 제목 키워드 기반

eXIAA: eXplainable Injections for Adversarial Attack

2025-11-13 · Leonardo Pesce, Jiawen Wei, Gianmarco Mengaldo arxiv

Post-hoc explainability methods are a subset of Machine Learning (ML) that aim to provide a reason for why a model behaves in a certain way. In this paper, we show a new black-box model-agnostic adversarial attack for po…

Adversarial Attack

Techniques for Adversarial Examples Threatening the Safety of Artificial Intelligence Based Systems

2019-09-29 · Utku Kose

Artificial intelligence is known as the most effective technological field for rapid developments shaping the future of the world. Even today, it is possible to see intense use of intelligence systems in all fields of th…

Boosting Adversarial Transferability Against Defenses via Multi-Scale Transformation

2025-07-02 · Zihong Guo, Chen Wan, Yayin Zheng, Hailing Kuang 외

The transferability of adversarial examples poses a significant security challenge for deep neural networks, which can be attacked without knowing anything about them. In this paper, we propose a new Segmented Gaussian P…

Post-Hoc Explanations Fail to Achieve their Purpose in Adversarial Contexts

2022-01-25 · Sebastian Bordt, Michèle Finck, Eric Raidl, Ulrike Von Luxburg

Existing and planned legislation stipulates various obligations to provide information about machine learning algorithms and their functioning, often interpreted as obligations to "explain". Many researchers suggest usin…

Exploring Counterfactual Explanations Through the Lens of Adversarial Examples: A Theoretical and Empirical Analysis

2021-06-18 · Martin Pawelczyk, Chirag Agarwal, Shalmali Joshi, Sohini Upadhyay 외

As machine learning (ML) models become more widely deployed in high-stakes applications, counterfactual explanations have emerged as key tools for providing actionable model explanations in practice. Despite the growing …

counterfactualCounterfactual Explanation