paper-with-me

홈 › Papers

Explaining Latent Representations with a Corpus of Examples

2021-10-28 · NeurIPS 2021 12 · Jonathan Crabbé, Zhaozhi Qian, Fergus Imrie, Mihaela van der Schaar

Modern machine learning models are complicated. Most of them rely on convoluted latent representations of their input to issue a prediction. To achieve greater transparency than a black-box that connects inputs to predictions, it is necessary to gain a deeper understanding of these latent representations. To that aim, we propose SimplEx: a user-centred method that provides example-based explanations with reference to a freely selected set of examples, called the corpus. SimplEx uses the corpus to improve the user's understanding of the latent space with post-hoc explanations answering two questions: (1) Which corpus examples explain the prediction issued for a given test example? (2) What features of these corpus examples are relevant for the model to relate them to the test example? SimplEx provides an answer by reconstructing the test latent representation as a mixture of corpus latent representations. Further, we propose a novel approach, the Integrated Jacobian, that allows SimplEx to make explicit the contribution of each corpus feature in the mixture. Through experiments on tasks ranging from mortality prediction to image classification, we demonstrate that these decompositions are robust and accurate. With illustrative use cases in medicine, we show that SimplEx empowers the user by highlighting relevant patterns in the corpus that explain model representations. Moreover, we demonstrate how the freedom in choosing the corpus allows the user to have personalized explanations in terms of examples that are meaningful for them.

📄 PDF Abstract BibTeX arXiv:2110.15355

Code (1)

jonathancrabbe/simplex 공식 구현 pytorch

Tasks

image-classificationImage ClassificationMortality Prediction

Methods 이 논문이 사용한 방법론

Test 설명 없음

Similar Papers 제목 키워드 기반

Explaining and Improving Model Behavior with k Nearest Neighbor Representations

2020-10-18 · Nazneen Fatema Rajani, Ben Krause, Wengpeng Yin, Tong Niu 외

Interpretability techniques in NLP have mainly focused on understanding individual predictions using attention visualization or gradient-based saliency maps over tokens. We propose using k nearest neighbor (kNN) represen…

Natural Language Inference

Beyond Examples: Constructing Explanation Space for Explaining Prototypes

2021-09-29 · Hyungjun Joo, Seokhyeon Ha, Jae Myung Kim, Sungyeob Han 외

As deep learning has been successfully deployed in diverse applications, there is ever increasing need for explaining its decision. Most of the existing methods produced explanations with a second model that explains the…

Contrastive Corpus Attribution for Explaining Representations

2022-09-30 · Chris Lin, Hugh Chen, Chanwoo Kim, Su-In Lee

Despite the widespread use of unsupervised models, very few methods are designed to explain them. Most explanation methods explain a scalar model output. However, unsupervised models output representation vectors, the el…

Contrastive LearningObject Localization

Explaining Classes through Word Attribution

2021-08-31 · Samuel Rönnqvist, Amanda Myntti, Aki-Juhani Kyröläinen, Sampo Pyysalo 외

In recent years, several methods have been proposed for explaining individual predictions of deep learning models, yet there has been little study of how to aggregate these predictions to explain how such models view cla…

ClassificationDeep LearningGenre classificationtext-classification+1

Topic-Partitioned Multinetwork Embeddings

2012-12-01 · NeurIPS 2012 12 · Peter Krafft, Juston Moore, Bruce Desmarais, Hanna M. Wallach

We introduce a joint model of network content and context designed for exploratory analysis of email networks via visualization of topic-specific communication patterns. Our model is an admixture model for text and netwo…

DescriptiveLink Prediction