Interpretation of Neural Networks is Fragile
In order for machine learning to be deployed and trusted in many applications, it is crucial to be able to reliably explain why the machine learning algorithm makes certain predictions. For example, if an algorithm classifies a given pathology image to be a malignant tumor, then the doctor may need to know which parts of the image led the algorithm to this classification. How to interpret black-box predictors is thus an important and active area of research. A fundamental question is: how much can we trust the interpretation itself? In this paper, we show that interpretation of deep learning predictions is extremely fragile in the following sense: two perceptively indistinguishable inputs with the same predicted label can be assigned very different interpretations. We systematically characterize the fragility of several widely-used feature-importance interpretation methods (saliency maps, relevance propagation, and DeepLIFT) on ImageNet and CIFAR-10. Our experiments show that even small random perturbation can change the feature importance and new systematic perturbations can lead to dramatically different interpretations without changing the label. We extend these results to show that interpretations based on exemplars (e.g. influence functions) are similarly fragile. Our analysis of the geometry of the Hessian matrix gives insight on why fragility could be a fundamental challenge to the current interpretation approaches.
Code (2)
Tasks
BIG-bench Machine LearningFeature ImportanceSimilar Papers 제목 키워드 기반
Identifying the Source of Vulnerability in Fragile Interpretations: A Case Study in Neural Text Classification
Prior works mainly used input perturbation methods for testing stability of post-hoc interpretation methods and observed fragile interpretations. However, different works show conflicting results on the primary source of…
text-classificationText ClassificationINTERPRETATION OF NEURAL NETWORK IS FRAGILE
In order for machine learning to be deployed and trusted in many applications, it is crucial to be able to reliably explain why the machine learning algorithm makes certain predictions. For example, if an algorithm class…
BIG-bench Machine LearningFeature ImportancePerturbing Inputs for Fragile Interpretations in Deep Natural Language Processing
Interpretability methods like Integrated Gradient and LIME are popular choices for explaining natural language model predictions with relative word importance scores. These interpretations need to be robust for trustwort…
Language ModelingLanguage ModellingUnlearning-based Neural Interpretations
Gradient-based interpretations often require an anchor point of comparison to avoid saturation in computing feature importance. We show that current baselines defined using static functions--constant mapping, averaging o…
Feature ImportanceLearning Gentle Grasping Using Vision, Sound, and Touch
In our daily life, we often encounter objects that are fragile and can be damaged by excessive grasping force, such as fruits. For these objects, it is paramount to grasp gently -- not using the maximum amount of force p…