paper-with-me

홈 › Papers

A Vulnerability of Attribution Methods Using Pre-Softmax Scores

2023-07-06 · Miguel Lerma, Mirtha Lucas

We discuss a vulnerability involving a category of attribution methods used to provide explanations for the outputs of convolutional neural networks working as classifiers. It is known that this type of networks are vulnerable to adversarial attacks, in which imperceptible perturbations of the input may alter the outputs of the model. In contrast, here we focus on effects that small modifications in the model may cause on the attribution method without altering the model outputs.

📄 PDF Abstract BibTeX arXiv:2307.03305

Code (1)

mlerma54/adversarial-attacks-on-saliency-maps 공식 구현

Methods 이 논문이 사용한 방법론

Focus 설명 없음

Similar Papers 제목 키워드 기반

Pre or Post-Softmax Scores in Gradient-based Attribution Methods, What is Best?

2023-06-22 · Miguel Lerma, Mirtha Lucas

Gradient based attribution methods for neural networks working as classifiers use gradients of network scores. Here we discuss the practical differences between using gradients of pre-softmax scores versus post-softmax s…

A Polynomial Architecture-Attribution Co-Design Framework for Exact Aumann-Shapley Attribution in GNNs

2026-07-23 · Bizu Feng, Zhimu Yang, Shuming Wang, Shaode Yu 외 arxiv

We study feature-level and node-level explanations for graph neural networks (GNNs) through the lens of Aumann-Shapley attribution. Path-integral methods such as Integrated Gradients provide an axiomatic formulation of a…

On Attribution of Recurrent Neural Network Predictions via Additive Decomposition

2019-03-27 · Mengnan Du, Ninghao Liu, Fan Yang, Shuiwang Ji 외

RNN models have achieved the state-of-the-art performance in a wide range of text mining tasks. However, these models are often regarded as black-boxes and are criticized due to the lack of interpretability. In this pape…

Decision Making

Improving Adversarial Robustness of Attribution via Implicit Regularization

2026-05-28 · Amir Mehrpanah, Matteo Gamba, Hossein Azizpour arxiv

The adversarial robustness of attributions is a fundamental requirement for reliable explainability in deep learning, yet existing approaches typically rely on computationally expensive explicit regularization. In this w…

Adversarial Robustness

SHIELD: Thwarting Code Authorship Attribution

2023-04-26 · Mohammed Abuhamad, Changhun Jung, David Mohaisen, DaeHun Nyang

Authorship attribution has become increasingly accurate, posing a serious privacy risk for programmers who wish to remain anonymous. In this paper, we introduce SHIELD to examine the robustness of different code authorsh…

Authorship Attribution