Semantics and explanation: why counterfactual explanations produce adversarial examples in deep neural networks
Recent papers in explainable AI have made a compelling case for counterfactual modes of explanation. While counterfactual explanations appear to be extremely effective in some instances, they are formally equivalent to adversarial examples. This presents an apparent paradox for explainability researchers: if these two procedures are formally equivalent, what accounts for the explanatory divide apparent between counterfactual explanations and adversarial examples? We resolve this paradox by placing emphasis back on the semantics of counterfactual expressions. Producing satisfactory explanations for deep learning systems will require that we find ways to interpret the semantics of hidden layer representations in deep neural networks.
Code (0)
등록된 구현이 없습니다.
Tasks
counterfactualSimilar Papers 제목 키워드 기반
ATEX-CF: Attack-Informed Counterfactual Explanations for Graph Neural Networks
Counterfactual explanations offer an intuitive way to interpret graph neural networks (GNNs) by identifying minimal changes that alter a model's prediction, thereby answering "what must differ for a different outcome?". …
Explanation GenerationNode ClassificationAdversarial AttackExploring Counterfactual Explanations Through the Lens of Adversarial Examples: A Theoretical and Empirical Analysis
As machine learning (ML) models become more widely deployed in high-stakes applications, counterfactual explanations have emerged as key tools for providing actionable model explanations in practice. Despite the growing …
counterfactualCounterfactual ExplanationSTEEX: Steering Counterfactual Explanations with Semantics
As deep learning models are increasingly used in safety-critical applications, explainability and trustworthiness become major concerns. For simple images, such as low-resolution face portraits, synthesizing visual count…
counterfactualCounterfactual ExplanationChoose your Data Wisely: A Framework for Semantic Counterfactuals
Counterfactual explanations have been argued to be one of the most intuitive forms of explanation. They are typically defined as a minimal set of edits on a given data sample that, when applied, changes the output of a m…
counterfactualCounterfactual ExplanationKnowledge GraphsCounterfactual Training: Teaching Models Plausible and Actionable Explanations
We propose a novel training regime termed counterfactual training that leverages counterfactual explanations to increase the explanatory capacity of models. Counterfactual explanations have emerged as a popular post-hoc …
Adversarial Robustness