paper-with-me

Papers

HIVE: Evaluating the Human Interpretability of Visual Explanations

2021-12-06 · Sunnie S. Y. Kim, Nicole Meister, Vikram V. Ramaswamy, Ruth Fong, Olga Russakovsky

As AI technology is increasingly applied to high-impact, high-risk domains, there have been a number of new methods aimed at making AI models more human interpretable. Despite the recent growth of interpretability work, there is a lack of systematic evaluation of proposed techniques. In this work, we introduce HIVE (Human Interpretability of Visual Explanations), a novel human evaluation framework that assesses the utility of explanations to human users in AI-assisted decision making scenarios, and enables falsifiable hypothesis testing, cross-method comparison, and human-centered evaluation of visual interpretability methods. To the best of our knowledge, this is the first work of its kind. Using HIVE, we conduct IRB-approved human studies with nearly 1000 participants and evaluate four methods that represent the diversity of computer vision interpretability works: GradCAM, BagNet, ProtoPNet, and ProtoTree. Our results suggest that explanations engender human trust, even for incorrect predictions, yet are not distinct enough for users to distinguish between correct and incorrect predictions. We open-source HIVE to enable future studies and encourage more human-centered approaches to interpretability research.

📄 PDF Abstract BibTeX arXiv:2112.03184

Code (1)

princetonvisualai/HIVE 공식 구현

Tasks

Decision MakingDiversity

Similar Papers 제목 키워드 기반

Would this change your answer? Evaluating Explanations of LLM Behavior In The Wild with Counterfactual Experiments

2026-08-17 · Adam Karvonen, Euan Ong, Subhash Kantamneni, Samuel Marks arxiv

Many areas of AI research, such as language model interpretability and chain of thought faithfulness, seek to explain model behaviors. But what constitutes a "good" explanation? In this work, we evaluate explanations thr…

A Study on Multimodal and Interactive Explanations for Visual Question Answering

2020-03-01 · Kamran Alipour, Jurgen P. Schulze, Yi Yao, Avi Ziskind 외

Explainability and interpretability of AI models is an essential factor affecting the safety of AI. While various explainable AI (XAI) approaches aim at mitigating the lack of transparency in deep networks, the evidence …

Explainable Artificial Intelligence (XAI)PredictionQuestion AnsweringVisual Question Answering+1

QUACKIE: A NLP Classification Task With Ground Truth Explanations

2020-12-24 · Yves Rychener, Xavier Renard, Djamé Seddah, Pascal Frossard 외

NLP Interpretability aims to increase trust in model predictions. This makes evaluating interpretability approaches a pressing issue. There are multiple datasets for evaluating NLP Interpretability, but their dependence …

ClassificationGeneral ClassificationQuestion Answering

Pitfalls in Evaluating Interpretability Agents

2026-03-20 · Tal Haklay, Nikhil Prakash, Sana Pandey, Antonio Torralba 외 arxiv

Automated interpretability systems aim to reduce the need for human labor and scale analysis to increasingly large models and diverse tasks. Recent efforts toward this goal leverage large language models (LLMs) at increa…

Evaluating SAE interpretability without explanations

2025-07-11 · Gonçalo Paulo, Nora Belrose arxiv

Sparse autoencoders (SAEs) and transcoders have become important tools for machine learning interpretability. However, measuring how interpretable they are remains challenging, with weak consensus about which benchmarks …

Explanation Generation