paper-with-me

Papers

A Meta Survey of Quality Evaluation Criteria in Explanation Methods

2022-03-25 · Helena Löfström, Karl Hammar, Ulf Johansson

Explanation methods and their evaluation have become a significant issue in explainable artificial intelligence (XAI) due to the recent surge of opaque AI models in decision support systems (DSS). Since the most accurate AI models are opaque with low transparency and comprehensibility, explanations are essential for bias detection and control of uncertainty. There are a plethora of criteria to choose from when evaluating explanation method quality. However, since existing criteria focus on evaluating single explanation methods, it is not obvious how to compare the quality of different methods. This lack of consensus creates a critical shortage of rigour in the field, although little is written about comparative evaluations of explanation methods. In this paper, we have conducted a semi-systematic meta-survey over fifteen literature surveys covering the evaluation of explainability to identify existing criteria usable for comparative evaluations of explanation methods. The main contribution in the paper is the suggestion to use appropriate trust as a criterion to measure the outcome of the subjective evaluation criteria and consequently make comparative evaluations possible. We also present a model of explanation quality aspects. In the model, criteria with similar definitions are grouped and related to three identified aspects of quality; model, explanation, and user. We also notice four commonly accepted criteria (groups) in the literature, covering all aspects of explanation quality: Performance, appropriate trust, explanation satisfaction, and fidelity. We suggest the model be used as a chart for comparative evaluations to create more generalisable research in explanation quality.

📄 PDF Abstract BibTeX arXiv:2203.13929

Code (0)

등록된 구현이 없습니다.

Tasks

Bias DetectionExplainable artificial intelligenceExplainable Artificial Intelligence (XAI)Survey

Similar Papers 제목 키워드 기반

Evaluating Step-by-step Reasoning Traces: A Survey

2025-02-17 · Jinu Lee, Julia Hockenmaier

Step-by-step reasoning is widely used to enhance the reasoning ability of large language models (LLMs) in complex problems. Evaluating the quality of reasoning traces is crucial for understanding and improving LLM reason…

Survey

APPLS: Evaluating Evaluation Metrics for Plain Language Summarization

2023-05-23 · Yue Guo, Tal August, Gondy Leroy, Trevor Cohen 외

While there has been significant development of models for Plain Language Summarization (PLS), evaluation remains a challenge. PLS lacks a dedicated assessment metric, and the suitability of text generation evaluation me…

InformativenessLanguage ModellingText GenerationText Simplification

DeepSurvey-Bench: Evaluating Academic Value of Automatically Generated Scientific Surveys

2026-01-13 · Guo-Biao Zhang, Ding-Yuan Liu, Da-Yi Wu, Tian Lan 외 arxiv

The rapid development of automated survey generation technology has made it increasingly important to establish a comprehensive benchmark to evaluate the quality of generated surveys. Most existing benchmarks first const…

Evaluating Explainability: A Framework for Systematic Assessment and Reporting of Explainable AI Features

2025-06-16 · Miguel A. Lago, Ghada Zamzmi, Brandon Eich, Jana G. Delfino

Explainability features are intended to provide insight into the internal mechanisms of an AI device, but there is a lack of evaluation techniques for assessing the quality of provided explanations. We propose a framewor…

Altruist: Argumentative Explanations through Local Interpretations of Predictive Models

2020-10-15 · Ioannis Mollas, Nick Bassiliades, Grigorios Tsoumakas

Explainable AI is an emerging field providing solutions for acquiring insights into automated systems' rationale. It has been put on the AI map by suggesting ways to tackle key ethical and societal issues. Existing expla…

BIG-bench Machine LearningFeature ImportanceInterpretable Machine Learning