paper-with-me

홈 › Papers

Towards a Theory of Faithfulness: Faithful Explanations of Differentiable Classifiers over Continuous Data

2022-05-19 · Nico Potyka, Xiang Yin, Francesca Toni

There is broad agreement in the literature that explanation methods should be faithful to the model that they explain, but faithfulness remains a rather vague term. We revisit faithfulness in the context of continuous data and propose two formal definitions of faithfulness for feature attribution methods. Qualitative faithfulness demands that scores reflect the true qualitative effect (positive vs. negative) of the feature on the model and quanitative faithfulness that the magnitude of scores reflect the true quantitative effect. We discuss under which conditions these requirements can be satisfied to which extent (local vs global). As an application of the conceptual idea, we look at differentiable classifiers over continuous data and characterize Gradient-scores as follows: every qualitatively faithful feature attribution method is qualitatively equivalent to Gradient-scores. Furthermore, if an attribution method is quantitatively faithful in the sense that changes of the output of the classifier are proportional to the scores of features, then it is either equivalent to gradient-scoring or it is based on an inferior approximation of the classifier. To illustrate the practical relevance of the theory, we experimentally demonstrate that popular attribution methods can fail to give faithful explanations in the setting where the data is continuous and the classifier differentiable.

📄 PDF Abstract BibTeX arXiv:2205.09620

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Explanation-based Training with Differentiable Insertion/Deletion Metric-aware Regularizers

2023-10-19 · Yuya Yoshikawa, Tomoharu Iwata

The quality of explanations for the predictions made by complex machine learning predictors is often measured using insertion and deletion metrics, which assess the faithfulness of the explanations, i.e., how accurately …

Zero-Shot Faithful Textual Explanations via Directional-Derivative Influence on Predictions

2026-05-16 · Toshinori Yamauchi, Hiroshi Kera, Kazuhiko Kawamoto arxiv

Zero-shot textual explanations aim to make image classifiers more transparent by probing their internal representations, without relying on task-specific supervision or LVLMs. However, existing methods often miss the fea…

CAuSE: Decoding Multimodal Classifiers using Faithful Natural Language Explanation

2025-12-07 · Dibyanayan Bandyopadhyay, Soham Bhattacharjee, Mohammed Hasanuzzaman, Asif Ekbal arxiv

Multimodal classifiers function as opaque black box models. While several techniques exist to interpret their predictions, very few of them are as intuitive and accessible as natural language explanations (NLEs). To buil…

Decision Making

Towards Faithful Natural Language Explanations: A Study Using Activation Patching in Large Language Models

2024-10-18 · Wei Jie Yeo, Ranjan Satapathy, Erik Cambria

Large Language Models (LLMs) are capable of generating persuasive Natural Language Explanations (NLEs) to justify their answers. However, the faithfulness of these explanations should not be readily trusted at face value…

Faithfulness vs. Plausibility: On the (Un)Reliability of Explanations from Large Language Models

2024-02-07 · Chirag Agarwal, Sree Harsha Tanneru, Himabindu Lakkaraju

Large Language Models (LLMs) are deployed as powerful tools for several natural language processing (NLP) applications. Recent works show that modern LLMs can generate self-explanations (SEs), which elicit their intermed…

Decision Making