paper-with-me

홈 › Papers

Towards Unifying Feature Attribution and Counterfactual Explanations: Different Means to the Same End

2020-11-10 · Ramaravind Kommiya Mothilal, Divyat Mahajan, Chenhao Tan, Amit Sharma

Feature attributions and counterfactual explanations are popular approaches to explain a ML model. The former assigns an importance score to each input feature, while the latter provides input examples with minimal changes to alter the model's predictions. To unify these approaches, we provide an interpretation based on the actual causality framework and present two key results in terms of their use. First, we present a method to generate feature attribution explanations from a set of counterfactual examples. These feature attributions convey how important a feature is to changing the classification outcome of a model, especially on whether a subset of features is necessary and/or sufficient for that change, which attribution-based methods are unable to provide. Second, we show how counterfactual examples can be used to evaluate the goodness of an attribution-based explanation in terms of its necessity and sufficiency. As a result, we highlight the complementarity of these two approaches. Our evaluation on three benchmark datasets - Adult-Income, LendingClub, and German-Credit - confirms the complementarity. Feature attribution methods like LIME and SHAP and counterfactual explanation methods like Wachter et al. and DiCE often do not agree on feature importance rankings. In addition, by restricting the features that can be modified for generating counterfactual examples, we find that the top-k features from LIME or SHAP are often neither necessary nor sufficient explanations of a model's prediction. Finally, we present a case study of different explanation methods on a real-world hospital triage problem

📄 PDF Abstract BibTeX arXiv:2011.04917

Code (1)

interpretml/DiCE tf

Tasks

Causal InferencecounterfactualCounterfactual ExplanationFeature Importance

Methods 이 논문이 사용한 방법론

SHAP 설명 없음
LIME LIME, or Local Interpretable Model-Agnostic Explanations, is an algorithm that can explain the predictions of any classifier or regressor in a faithful way, by…
Causal inference Causal inference is the process of drawing a conclusion about a causal connection based on the conditions of the occurrence of an effect. The main difference between causal…

Similar Papers 제목 키워드 기반

On the Connection between Game-Theoretic Feature Attributions and Counterfactual Explanations

2023-07-13 · Emanuele Albini, Shubham Sharma, Saumitra Mishra, Danial Dervovic 외

Explainable Artificial Intelligence (XAI) has received widespread interest in recent years, and two of the most popular types of explanations are feature attributions, and counterfactual explanations. These classes of ap…

counterfactualCounterfactual ExplanationExplainable artificial intelligenceExplainable Artificial Intelligence (XAI)+1

Not All Explanations Simulate Equally: Comparing Verbalized Feature Attributions and Self-Generated Rationales

2026-05-31 · Pingjun Hong, Benjamin Roth arxiv

Natural-language explanations are often treated as a unified interface for understanding model behavior, but different explanation sources may support simulation in different ways. This paper compares two families of exp…

Question Answering

Counterfactual Shapley Additive Explanations

2021-10-27 · Emanuele Albini, Jason Long, Danial Dervovic, Daniele Magazzeni

Feature attributions are a common paradigm for model explanations due to their simplicity in assigning a single numeric score for each input feature to a model. In the actionable recourse setting, wherein the goal of the…

counterfactualCounterfactual ExplanationExplainable artificial intelligenceFeature Importance

Motif-guided Time Series Counterfactual Explanations

2022-11-08 · Peiyu Li, Soukaina Filali Boubrahimi, Shah Muhammad Hamdi

With the rising need of interpretable machine learning methods, there is a necessity for a rise in human effort to provide diverse explanations of the influencing factors of the model decisions. To improve the trust and …

counterfactualCounterfactual ExplanationDecision MakingExplainable artificial intelligence+5

Attribution-Scores and Causal Counterfactuals as Explanations in Artificial Intelligence

2023-03-06 · Leopoldo Bertossi

In this expository article we highlight the relevance of explanations for artificial intelligence, in general, and for the newer developments in {\em explainable AI}, referring to origins and connections of and among dif…

Logical ReasoningManagement