paper-with-me

홈 › Papers

An Evaluation of the Human-Interpretability of Explanation

2019-01-31 · Isaac Lage, Emily Chen, Jeffrey He, Menaka Narayanan, Been Kim, Sam Gershman, Finale Doshi-Velez

Recent years have seen a boom in interest in machine learning systems that can provide a human-understandable rationale for their predictions or decisions. However, exactly what kinds of explanation are truly human-interpretable remains poorly understood. This work advances our understanding of what makes explanations interpretable under three specific tasks that users may perform with machine learning systems: simulation of the response, verification of a suggested response, and determining whether the correctness of a suggested response changes under a change to the inputs. Through carefully controlled human-subject experiments, we identify regularizers that can be used to optimize for the interpretability of machine learning systems. Our results show that the type of complexity matters: cognitive chunks (newly defined concepts) affect performance more than variable repetitions, and these trends are consistent across tasks and domains. This suggests that there may exist some common design principles for explanation systems.

📄 PDF Abstract BibTeX arXiv:1902.00006

Code (0)

등록된 구현이 없습니다.

Tasks

BIG-bench Machine Learning

Methods 이 논문이 사용한 방법론

Interpretability 설명 없음

Similar Papers 제목 키워드 기반

HIVE: Evaluating the Human Interpretability of Visual Explanations

2021-12-06 · Sunnie S. Y. Kim, Nicole Meister, Vikram V. Ramaswamy, Ruth Fong 외

As AI technology is increasingly applied to high-impact, high-risk domains, there have been a number of new methods aimed at making AI models more human interpretable. Despite the recent growth of interpretability work, …

Decision MakingDiversity

Evaluating SAE interpretability without explanations

2025-07-11 · Gonçalo Paulo, Nora Belrose arxiv

Sparse autoencoders (SAEs) and transcoders have become important tools for machine learning interpretability. However, measuring how interpretable they are remains challenging, with weak consensus about which benchmarks …

Explanation Generation

The Promise and Peril of Human Evaluation for Model Interpretability

2017-11-20 · Bernease Herman

Transparency, user trust, and human comprehension are popular ethical motivations for interpretable machine learning. In support of these goals, researchers evaluate model explanation performance using humans and real wo…

DescriptiveInterpretable Machine LearningPosition

Concept-based Explanations using Non-negative Concept Activation Vectors and Decision Tree for CNN Models

2022-11-19 · Gayda Mutahar, Tim Miller

This paper evaluates whether training a decision tree based on concepts extracted from a concept-based explainer can increase interpretability for Convolutional Neural Networks (CNNs) models and boost the fidelity and pe…

From Human Explanation to Model Interpretability: A Framework Based on Weight of Evidence

2021-04-27 · David Alvarez-Melis, Harmanpreet Kaur, Hal Daumé III, Hanna Wallach 외

We take inspiration from the study of human explanation to inform the design and evaluation of interpretability methods in machine learning. First, we survey the literature on human explanation in philosophy, cognitive s…

BIG-bench Machine LearningInterpretable Machine LearningPhilosophy