paper-with-me

홈 › Papers

Interpretability is in the Mind of the Beholder: A Causal Framework for Human-interpretable Representation Learning

2023-09-14 · Emanuele Marconato, Andrea Passerini, Stefano Teso

Focus in Explainable AI is shifting from explanations defined in terms of low-level elements, such as input features, to explanations encoded in terms of interpretable concepts learned from data. How to reliably acquire such concepts is, however, still fundamentally unclear. An agreed-upon notion of concept interpretability is missing, with the result that concepts used by both post-hoc explainers and concept-based neural networks are acquired through a variety of mutually incompatible strategies. Critically, most of these neglect the human side of the problem: a representation is understandable only insofar as it can be understood by the human at the receiving end. The key challenge in Human-interpretable Representation Learning (HRL) is how to model and operationalize this human element. In this work, we propose a mathematical framework for acquiring interpretable representations suitable for both post-hoc explainers and concept-based neural networks. Our formalization of HRL builds on recent advances in causal representation learning and explicitly models a human stakeholder as an external observer. This allows us to derive a principled notion of alignment between the machine representation and the vocabulary of concepts understood by the human. In doing so, we link alignment and interpretability through a simple and intuitive name transfer game, and clarify the relationship between alignment and a well-known property of representations, namely disentanglment. We also show that alignment is linked to the issue of undesirable correlations among concepts, also known as concept leakage, and to content-style separation, all through a general information-theoretic reformulation of these properties. Our conceptualization aims to bridge the gap between the human and algorithmic sides of interpretability and establish a stepping stone for new research on human-interpretable representations.

📄 PDF Abstract BibTeX arXiv:2309.07742

Code (0)

등록된 구현이 없습니다.

Tasks

Representation Learning

Similar Papers 제목 키워드 기반

Deep Learning Predicts Biomarker Status and Discovers Related Histomorphology Characteristics for Low-Grade Glioma

2023-10-11 · Zijie Fang, Yihan Liu, Yifeng Wang, Xiangyang Zhang 외

Biomarker detection is an indispensable part in the diagnosis and treatment of low-grade glioma (LGG). However, current LGG biomarker detection methods rely on expensive and complex molecular genetic testing, for which p…

Multiple Instance LearningOne-Class Classificationwhole slide images

In the Eye of the Beholder: Robust Prediction with Causal User Modeling

2022-06-01 · Amir Feder, Guy Horowitz, Yoav Wald, Roi Reichart 외

Accurately predicting the relevance of items to users is crucial to the success of many social platforms. Conventional approaches train models on logged historical data; but recommendation systems, media services, and on…

Recommendation Systems

Causality and deceit: Do androids watch action movies?

2019-10-10 · Dusko Pavlovic, Temra Pavlovic

We seek causes through science, religion, and in everyday life. We get excited when a big rock causes a big splash, and we get scared when it tumbles without a cause. But our causal cognition is usually biased. The 'why'…

Causal potency of consciousness in the physical world

2023-06-26 · Danko D. Georgiev

The evolution of the human mind through natural selection mandates that our conscious experiences are causally potent in order to leave a tangible impact upon the surrounding physical world. Any attempt to construct a fu…

Beholder-GAN: Generation and Beautification of Facial Images with Conditioning on Their Beauty Level

2019-02-07 · Nir Diamant, Dean Zadok, Chaim Baskin, Eli Schwartz 외

Beauty is in the eye of the beholder. This maxim, emphasizing the subjectivity of the perception of beauty, has enjoyed a wide consensus since ancient times. In the digitalera, data-driven methods have been shown to be a…