paper-with-me

홈 › Papers

Interpreting Interpretations: Organizing Attribution Methods by Criteria

2020-02-19 · Zifan Wang, Piotr Mardziel, Anupam Datta, Matt Fredrikson

Motivated by distinct, though related, criteria, a growing number of attribution methods have been developed tointerprete deep learning. While each relies on the interpretability of the concept of "importance" and our ability to visualize patterns, explanations produced by the methods often differ. As a result, input attribution for vision models fail to provide any level of human understanding of model behaviour. In this work we expand the foundationsof human-understandable concepts with which attributionscan be interpreted beyond "importance" and its visualization; we incorporate the logical concepts of necessity andsufficiency, and the concept of proportionality. We definemetrics to represent these concepts as quantitative aspectsof an attribution. This allows us to compare attributionsproduced by different methods and interpret them in novelways: to what extent does this attribution (or this method)represent the necessity or sufficiency of the highlighted inputs, and to what extent is it proportional? We evaluate our measures on a collection of methods explaining convolutional neural networks (CNN) for image classification. We conclude that some attribution methods are more appropriate for interpretation in terms of necessity while others are in terms of sufficiency, while no method is always the most appropriate in terms of both.

📄 PDF Abstract BibTeX arXiv:2002.07985

Code (0)

등록된 구현이 없습니다.

Tasks

image-classificationImage Classification

Methods 이 논문이 사용한 방법론

Interpretability 설명 없음

Similar Papers 제목 키워드 기반

A Robust Unsupervised Ensemble of Feature-Based Explanations using Restricted Boltzmann Machines

2021-11-14 · Vadim Borisov, Johannes Meier, Johan van den Heuvel, Hamed Jalali 외

Understanding the results of deep neural networks is an essential step towards wider acceptance of deep learning algorithms. Many approaches address the issue of interpreting artificial neural networks, but often provide…

Interpreting Deep Neural Networks with the Package innsight

2023-06-19 · Niklas Koenen, Marvin N. Wright

The R package innsight offers a general toolbox for revealing variable-wise interpretations of deep neural networks' predictions with so-called feature attribution methods. Aside from the unified and user-friendly framew…

Influence Tuning: Demoting Spurious Correlations via Instance Attribution and Instance-Driven Updates

2021-10-07 · Findings (EMNLP) 2021 11 · Xiaochuang Han, Yulia Tsvetkov

Among the most critical limitations of deep learning NLP models are their lack of interpretability, and their reliance on spurious correlations. Prior work proposed various approaches to interpreting the black-box models…

Logic Traps in Evaluating Attribution Scores

2021-11-16 · ACL ARR November 2021 11 · Anonymous

Modern deep learning models are notoriously opaque, which has motivated the development of methods for interpreting how deep models predict. This goal is usually approached with attribution method, which assesses the inf…

Logic Traps in Evaluating Attribution Scores

2021-09-12 · ACL 2022 5 · Yiming Ju, Yuanzhe Zhang, Zhao Yang, Zhongtao Jiang 외

Modern deep learning models are notoriously opaque, which has motivated the development of methods for interpreting how deep models predict. This goal is usually approached with attribution method, which assesses the inf…