Global Explanations of Neural Networks: Mapping the Landscape of Predictions
A barrier to the wider adoption of neural networks is their lack of interpretability. While local explanation methods exist for one prediction, most global attributions still reduce neural network decisions to a single set of features. In response, we present an approach for generating global attributions called GAM, which explains the landscape of neural network predictions across subpopulations. GAM augments global explanations with the proportion of samples that each attribution best explains and specifies which samples are described by each attribution. Global explanations also have tunable granularity to detect more or fewer subpopulations. We demonstrate that GAM's global explanations 1) yield the known feature importances of simulated data, 2) match feature weights of interpretable statistical models on real data, and 3) are intuitive to practitioners through user studies. With more transparent predictions, GAM can help ensure neural network decisions are generated for the right reasons.
Code (1)
Methods 이 논문이 사용한 방법론
Similar Papers 제목 키워드 기반
Global Graph Counterfactual Explanation: A Subgraph Mapping Approach
Graph Neural Networks (GNNs) have been widely deployed in various real-world applications. However, most GNNs are black-box models that lack explanations. One strategy to explain GNNs is through counterfactual explanatio…
counterfactualCounterfactual ExplanationLAMP: Extracting Locally Linear Decision Surfaces from LLM World Models
We introduce LAMP (Linear Attribution Mapping Probe), a method that shines light onto a black-box language model's decision surface and studies how reliably a model maps its stated reasons to its predictions through a lo…
Sentiment AnalysisUtilizing Description Logics for Global Explanations of Heterogeneous Graph Neural Networks
Graph Neural Networks (GNNs) are effective for node classification in graph-structured data, but they lack explainability, especially at the global level. Current research mainly utilizes subgraphs of the input as local …
Node ClassificationDissenting Explanations: Leveraging Disagreement to Reduce Model Overreliance
While explainability is a desirable characteristic of increasingly complex black-box models, modern explanation methods have been shown to be inconsistent and contradictory. The semantics of explanations is not always fu…
modelLearning Global Transparent Models Consistent with Local Contrastive Explanations
There is a rich and growing literature on producing local contrastive/counterfactual explanations for black-box models (e.g. neural networks). In these methods, for an input, an explanation is in the form of a contrast p…
counterfactual