paper-with-me

홈 › Papers

Proxy Tasks and Subjective Measures Can Be Misleading in Evaluating Explainable AI Systems

2020-01-22 · Zana Buçinca, Phoebe Lin, Krzysztof Z. Gajos, Elena L. Glassman

Explainable artificially intelligent (XAI) systems form part of sociotechnical systems, e.g., human+AI teams tasked with making decisions. Yet, current XAI systems are rarely evaluated by measuring the performance of human+AI teams on actual decision-making tasks. We conducted two online experiments and one in-person think-aloud study to evaluate two currently common techniques for evaluating XAI systems: (1) using proxy, artificial tasks such as how well humans predict the AI's decision from the given explanations, and (2) using subjective measures of trust and preference as predictors of actual performance. The results of our experiments demonstrate that evaluations with proxy tasks did not predict the results of the evaluations with the actual decision-making tasks. Further, the subjective measures on evaluations with actual decision-making tasks did not predict the objective performance on those same tasks. Our results suggest that by employing misleading evaluation methods, our field may be inadvertently slowing its progress toward developing human+AI teams that can reliably perform better than humans or AIs alone.

📄 PDF Abstract BibTeX arXiv:2001.08298

Code (0)

등록된 구현이 없습니다.

Tasks

Decision MakingExplainable Artificial Intelligence (XAI)

Similar Papers 제목 키워드 기반

Evaluating Causal Models by Comparing Interventional Distributions

2016-08-16 · Dan Garant, David Jensen

The predominant method for evaluating the quality of causal models is to measure the graphical accuracy of the learned model structure. We present an alternative method for evaluating causal models that directly measures…

Study on the Correlation between Objective Evaluations and Subjective Speech Quality and Intelligibility

2023-07-10 · Hsin-Tien Chiang, Kuo-Hsuan Hung, Szu-Wei Fu, Heng-Cheng Kuo 외

Subjective tests are the gold standard for evaluating speech quality and intelligibility; however, they are time-consuming and expensive. Thus, objective measures that align with human perceptions are crucial. This study…

Deep Learning

PerSEval: Assessing Personalization in Text Summarizers

2024-06-29 · Sourish Dasgupta, Ankush Chander, Parth Borad, Isha Motiyani 외

Personalized summarization models cater to individuals' subjective understanding of saliency, as represented by their reading history and current topics of attention. Existing personalized text summarizers are primarily …

BenchmarkingHuman Judgment Correlation

Full Reference Objective Quality Assessment for Reconstructed Background Images

2018-03-12 · Aditee Shrotre, Lina Karam

With an increased interest in applications that require a clean background image, such as video surveillance, object tracking, street view imaging and location-based services on web-based maps, multiple algorithms have b…

Image Quality AssessmentObject Tracking

Extrinsically-Focused Evaluation of Omissions in Medical Summarization

2023-11-14 · Elliot Schumacher, Daniel Rosenthal, Dhruv Naik, Varun Nair 외

Large language models (LLMs) have shown promise in safety-critical applications such as healthcare, yet the ability to quantify performance has lagged. An example of this challenge is in evaluating a summary of the patie…

Decision Making