paper-with-me

홈 › Papers

Comparing Automatic and Human Evaluation of Local Explanations for Text Classification

2018-06-01 · NAACL 2018 6 · Dong Nguyen

Text classification models are becoming increasingly complex and opaque, however for many applications it is essential that the models are interpretable. Recently, a variety of approaches have been proposed for generating local explanations. While robust evaluations are needed to drive further progress, so far it is unclear which evaluation approaches are suitable. This paper is a first step towards more robust evaluations of local explanations. We evaluate a variety of local explanation approaches using automatic measures based on word deletion. Furthermore, we show that an evaluation using a crowdsourcing experiment correlates moderately with these automatic measures and that a variety of other factors also impact the human judgements.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Tasks

General ClassificationRecommendation Systemstext-classificationText Classification

Similar Papers 제목 키워드 기반

A Study of Automatic Metrics for the Evaluation of Natural Language Explanations

2021-03-15 · EACL 2021 2 · Miruna Clinciu, Arash Eshghi, Helen Hastie

As transparency becomes key for robotics and AI, it will be necessary to evaluate the methods through which transparency is provided, including automatically generated natural language (NL) explanations. Here, we explore…

nlg evaluationText Generation

A Human-Grounded Evaluation Benchmark for Local Explanations of Machine Learning

2018-01-16 · Sina Mohseni, Jeremy E. Block, Eric D. Ragan

Research in interpretable machine learning proposes different computational and human subject approaches to evaluate model saliency explanations. These approaches measure different qualities of explanations to achieve di…

BIG-bench Machine LearningDecision MakingInterpretable Machine LearningSegmentation+1

Evaluating Model Explanations without Ground Truth

2025-05-15 · Kaivalya Rawal, Zihao Fu, Eoin Delaney, Chris Russell

There can be many competing and contradictory explanations for a single model prediction, making it difficult to select which one to use. Current explanation evaluation frameworks measure quality by comparing against ide…

Feature ImportancemodelSensitivity

Interpreting Language Reward Models via Contrastive Explanations

2024-11-25 · Junqi Jiang, Tom Bewley, Saumitra Mishra, Freddy Lecue 외

Reward models (RMs) are a crucial component in the alignment of large language models' (LLMs) outputs with human values. RMs approximate human preferences over possible LLM responses to the same prompt by predicting and …

Attribute

FunnyBirds: A Synthetic Vision Dataset for a Part-Based Analysis of Explainable AI Methods

2023-08-11 · ICCV 2023 1 · Robin Hesse, Simone Schaub-Meyer, Stefan Roth

The field of explainable artificial intelligence (XAI) aims to uncover the inner workings of complex deep neural models. While being crucial for safety-critical domains, XAI inherently lacks ground-truth explanations, ma…

Explainable artificial intelligenceExplainable Artificial Intelligence (XAI)