paper-with-me

Papers

XEQ Scale for Evaluating XAI Experience Quality

2024-07-15 · Anjana Wijekoon, Nirmalie Wiratunga, David Corsar, Kyle Martin, Ikechukwu Nkisi-Orji, Belen Díaz-Agudo, Derek Bridge

Explainable Artificial Intelligence (XAI) aims to improve the transparency of autonomous decision-making through explanations. Recent literature has emphasised users' need for holistic "multi-shot" explanations and personalised engagement with XAI systems. We refer to this user-centred interaction as an XAI Experience. Despite advances in creating XAI experiences, evaluating them in a user-centred manner has remained challenging. In response, we developed the XAI Experience Quality (XEQ) Scale. XEQ quantifies the quality of experiences across four dimensions: learning, utility, fulfilment and engagement. These contributions extend the state-of-the-art of XAI evaluation, moving beyond the one-dimensional metrics frequently developed to assess single-shot explanations. This paper presents the XEQ scale development and validation process, including content validation with XAI experts, and discriminant and construct validation through a large-scale pilot study. Our pilot study results offer strong evidence that establishes the XEQ Scale as a comprehensive framework for evaluating user-centred XAI experiences.

📄 PDF Abstract BibTeX arXiv:2407.10662

Code (0)

등록된 구현이 없습니다.

Tasks

Decision MakingExplainable artificial intelligenceExplainable Artificial Intelligence (XAI)

Similar Papers 제목 키워드 기반

Evaluating Sensor Data Quality in Internet ofThings Smart Agriculture Applications

2021-04-28 · Kaneez Fizza, Prem Prakash Jayaraman, Abhik Banerjee, Dimitrios Georgakopoulos 외

The unprecedented growth of Internet of Things (IoT) and its applications in areas such as Smart Agriculture compels the need to devise newer ways for evaluating the quality of such applications. While existing models fo…

Understanding the Impact of Experiment Design for Evaluating Dialogue System Output

2020-07-01 · WS 2020 7 · Sashank Santhanam, Samira Shaikh

Evaluation of output from natural language generation (NLG) systems is typically conducted via crowdsourced human judgments. To understand the impact of how experiment design might affect the quality and consistency of s…

Text Generation

Towards Best Experiment Design for Evaluating Dialogue System Output

2019-09-23 · WS 2019 10 · Sashank Santhanam, Samira Shaikh

To overcome the limitations of automated metrics (e.g. BLEU, METEOR) for evaluating dialogue systems, researchers typically use human judgments to provide convergent evidence. While it has been demonstrated that human ju…

Dialogue Evaluation

Automated Code Review Using Large Language Models at Ericsson: An Experience Report

2025-07-25 · Shweta Ramesh, Joy Bose, Hamender Singh, A K Raghavan 외 arxiv

Code review is one of the primary means of assuring the quality of released software along with testing and static analysis. However, code review requires experienced developers who may not always have the time to perfor…

Optimizing and Evaluating Enterprise Retrieval-Augmented Generation (RAG): A Content Design Perspective

2024-10-01 · Sarah Packowski, Inge Halilovic, Jenifer Schlotfeldt, Trish Smith

Retrieval-augmented generation (RAG) is a popular technique for using large language models (LLMs) to build customer-support, question-answering solutions. In this paper, we share our team's practical experience building…

Question AnsweringRAGRetrievalRetrieval-augmented Generation