paper-with-me

홈 › Papers

The Intriguing Properties of Model Explanations

2018-01-30 · Maruan Al-Shedivat, Avinava Dubey, Eric P. Xing

Linear approximations to the decision boundary of a complex model have become one of the most popular tools for interpreting predictions. In this paper, we study such linear explanations produced either post-hoc by a few recent methods or generated along with predictions with contextual explanation networks (CENs). We focus on two questions: (i) whether linear explanations are always consistent or can be misleading, and (ii) when integrated into the prediction process, whether and how explanations affect the performance of the model. Our analysis sheds more light on certain properties of explanations produced by different methods and suggests that learning models that explain and predict jointly is often advantageous.

📄 PDF Abstract BibTeX arXiv:1801.09808

Code (1)

alshedivat/cen 공식 구현 tf

Tasks

model

Similar Papers 제목 키워드 기반

What Gets Echoed? Understanding the "Pointers" in Explanations of Persuasive Arguments

2019-11-01 · David Atkinson, Kumar Bhargav Srinivasan, Chenhao Tan

Explanations are central to everyday life, and are a topic of growing interest in the AI community. To investigate the process of providing natural language explanations, we leverage the dynamics of the /r/ChangeMyView s…

What Gets Echoed? Understanding the ``Pointers'' in Explanations of Persuasive Arguments

2019-11-01 · IJCNLP 2019 11 · David Atkinson, Kumar Bhargav Srinivasan, Chenhao Tan

Explanations are central to everyday life, and are a topic of growing interest in the AI community. To investigate the process of providing natural language explanations, we leverage the dynamics of the /r/ChangeMyView s…

The Intriguing Relation Between Counterfactual Explanations and Adversarial Examples

2020-09-11 · Timo Freiesleben

The same method that creates adversarial examples (AEs) to fool image-classifiers can be used to generate counterfactual explanations (CEs) that explain algorithmic decisions. This observation has led researchers to cons…

counterfactualRelation

Explaining Machine Learning Classifiers through Diverse Counterfactual Explanations

2019-05-19 · Ramaravind Kommiya Mothilal, Amit Sharma, Chenhao Tan

Post-hoc explanations of machine learning models are crucial for people to understand and act on algorithmic predictions. An intriguing class of explanations is through counterfactuals, hypothetical examples that show pe…

BIG-bench Machine LearningcounterfactualDiversityPoint Processes

Explaining Away Attacks Against Neural Networks

2020-03-06 · Sean Saito, Jin Wang

We investigate the problem of identifying adversarial attacks on image-based neural networks. We present intriguing experimental results showing significant discrepancies between the explanations generated for the predic…