paper-with-me

Papers

Exploring Counterfactual Explanations Through the Lens of Adversarial Examples: A Theoretical and Empirical Analysis

2021-06-18 · Martin Pawelczyk, Chirag Agarwal, Shalmali Joshi, Sohini Upadhyay, Himabindu Lakkaraju

As machine learning (ML) models become more widely deployed in high-stakes applications, counterfactual explanations have emerged as key tools for providing actionable model explanations in practice. Despite the growing popularity of counterfactual explanations, a deeper understanding of these explanations is still lacking. In this work, we systematically analyze counterfactual explanations through the lens of adversarial examples. We do so by formalizing the similarities between popular counterfactual explanation and adversarial example generation methods identifying conditions when they are equivalent. We then derive the upper bounds on the distances between the solutions output by counterfactual explanation and adversarial example generation methods, which we validate on several real-world data sets. By establishing these theoretical and empirical similarities between counterfactual explanations and adversarial examples, our work raises fundamental questions about the design and development of existing counterfactual explanation algorithms.

📄 PDF Abstract BibTeX arXiv:2106.09992

Code (0)

등록된 구현이 없습니다.

Tasks

counterfactualCounterfactual Explanation

Similar Papers 제목 키워드 기반

Generating Plausible Counterfactual Explanations for Deep Transformers in Financial Text Classification

2020-10-23 · COLING 2020 8 · Linyi Yang, Eoin M. Kenny, Tin Lok James Ng, Yi Yang 외

Corporate mergers and acquisitions (M&A) account for billions of dollars of investment globally every year, and offer an interesting and challenging domain for artificial intelligence. However, in these highly sensitive …

counterfactualExplainable Artificial Intelligence (XAI)General Classificationtext-classification+1

Would this change your answer? Evaluating Explanations of LLM Behavior In The Wild with Counterfactual Experiments

2026-08-17 · Adam Karvonen, Euan Ong, Subhash Kantamneni, Samuel Marks arxiv

Many areas of AI research, such as language model interpretability and chain of thought faithfulness, seek to explain model behaviors. But what constitutes a "good" explanation? In this work, we evaluate explanations thr…

Semantics and explanation: why counterfactual explanations produce adversarial examples in deep neural networks

2020-12-18 · Kieran Browne, Ben Swift

Recent papers in explainable AI have made a compelling case for counterfactual modes of explanation. While counterfactual explanations appear to be extremely effective in some instances, they are formally equivalent to a…

counterfactual

Counterfactual Explanations for Graph Classification Through the Lenses of Density

2023-07-27 · Carlo Abrate, Giulia Preti, Francesco Bonchi

Counterfactual examples have emerged as an effective approach to produce simple and understandable post-hoc explanations. In the context of graph classification, previous work has focused on generating counterfactual exp…

counterfactualCounterfactual ExplanationGraph Classification

Adversarial Counterfactual Visual Explanations

2023-03-17 · CVPR 2023 1 · Guillaume Jeanneret, Loïc Simon, Frédéric Jurie

Counterfactual explanations and adversarial attacks have a related goal: flipping output labels with minimal perturbations regardless of their characteristics. Yet, adversarial attacks cannot be used directly in a counte…

counterfactualCounterfactual ExplanationDenoising