paper-with-me

홈 › Papers

Discriminative Attribution from Counterfactuals

2021-09-28 · Nils Eckstein, Alexander S. Bates, Gregory S. X. E. Jefferis, Jan Funke

We present a method for neural network interpretability by combining feature attribution with counterfactual explanations to generate attribution maps that highlight the most discriminative features between pairs of classes. We show that this method can be used to quantitatively evaluate the performance of feature attribution methods in an objective manner, thus preventing potential observer bias. We evaluate the proposed method on three diverse datasets, including a challenging artificial dataset and real-world biological data. We show quantitatively and qualitatively that the highlighted features are substantially more discriminative than those extracted using conventional attribution methods and argue that this type of explanation is better suited for understanding fine grained class differences as learned by a deep neural network.

📄 PDF Abstract BibTeX arXiv:2109.13412

Code (0)

등록된 구현이 없습니다.

Tasks

counterfactual

Similar Papers 제목 키워드 기반

Connecting Attributions and QA Model Behavior on Realistic Counterfactuals

2021-04-09 · EMNLP 2021 11 · Xi Ye, Rohan Nair, Greg Durrett

When a model attribution technique highlights a particular part of the input, a user might understand this highlight as making a statement about counterfactuals (Miller, 2019): if that part of the input were to change, t…

counterfactualMachine Reading ComprehensionReading ComprehensionSentiment Analysis

Attribution-Scores and Causal Counterfactuals as Explanations in Artificial Intelligence

2023-03-06 · Leopoldo Bertossi

In this expository article we highlight the relevance of explanations for artificial intelligence, in general, and for the newer developments in {\em explainable AI}, referring to origins and connections of and among dif…

Logical ReasoningManagement

FitCF: A Framework for Automatic Feature Importance-guided Counterfactual Example Generation

2025-01-01 · Qianli Wang, Nils Feldhus, Simon Ostermann, Luis Felipe Villa-Arenas 외

Counterfactual examples are widely used in natural language processing (NLP) as valuable data to improve models, and in explainable artificial intelligence (XAI) to understand model behavior. The automated generation of …

counterfactualExplainable artificial intelligenceExplainable Artificial Intelligence (XAI)Feature Importance

Understanding and evaluating computer vision models through the lens of counterfactuals

2025-08-28 · Pushkar Shukla arxiv

Counterfactual reasoning -- the practice of asking ``what if'' by varying inputs and observing changes in model behavior -- has become central to interpretable and fair AI. This thesis develops frameworks that use counte…

Image Generation

Counterfactuals As a Means for Evaluating Faithfulness of Attribution Methods in Autoregressive Language Models

2024-08-21 · Sepehr Kamahi, Yadollah Yaghoobzadeh

Despite the widespread adoption of autoregressive language models, explainability evaluation research has predominantly focused on span infilling and masked language models. Evaluating the faithfulness of an explanation …

counterfactualDecision MakingFeature ImportanceLanguage Modelling