paper-with-me

홈 › Papers

Human Evaluation of Spoken vs. Visual Explanations for Open-Domain QA

2020-12-30 · Ana Valeria Gonzalez, Gagan Bansal, Angela Fan, Robin Jia, Yashar Mehdad, Srinivasan Iyer

While research on explaining predictions of open-domain QA systems (ODQA) to users is gaining momentum, most works have failed to evaluate the extent to which explanations improve user trust. While few works evaluate explanations using user studies, they employ settings that may deviate from the end-user's usage in-the-wild: ODQA is most ubiquitous in voice-assistants, yet current research only evaluates explanations using a visual display, and may erroneously extrapolate conclusions about the most performant explanations to other modalities. To alleviate these issues, we conduct user studies that measure whether explanations help users correctly decide when to accept or reject an ODQA system's answer. Unlike prior work, we control for explanation modality, e.g., whether they are communicated to users through a spoken or visual interface, and contrast effectiveness across modalities. Our results show that explanations derived from retrieved evidence passages can outperform strong baselines (calibrated confidence) across modalities but the best explanation strategy in fact changes with the modality. We show common failure cases of current explanations, emphasize end-to-end evaluation of explanations, and caution against evaluating them in proxy modalities that are different from deployment.

📄 PDF Abstract BibTeX arXiv:2012.15075

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Do Explanations Help Users Detect Errors in Open-Domain QA? An Evaluation of Spoken vs. Visual Explanations

2021-08-01 · Findings (ACL) 2021 8 · Ana Valeria González, Gagan Bansal, Angela Fan, Yashar Mehdad 외

AudioMNIST: Exploring Explainable Artificial Intelligence for Audio Analysis on a Simple Benchmark

2018-07-09 · Sören Becker, Johanna Vielhaben, Marcel Ackermann, Klaus-Robert Müller 외

Explainable Artificial Intelligence (XAI) is targeted at understanding how models perform feature selection and derive their classification decisions. This paper explores post-hoc explanations for deep neural networks in…

Audio ClassificationDecision MakingExplainable artificial intelligenceExplainable Artificial Intelligence (XAI)+2

HIVE: Evaluating the Human Interpretability of Visual Explanations

2021-12-06 · Sunnie S. Y. Kim, Nicole Meister, Vikram V. Ramaswamy, Ruth Fong 외

As AI technology is increasingly applied to high-impact, high-risk domains, there have been a number of new methods aimed at making AI models more human interpretable. Despite the recent growth of interpretability work, …

Decision MakingDiversity

Graphical Perception of Saliency-based Model Explanations

2024-06-11 · Yayan Zhao, MingWei Li, Matthew Berger

In recent years, considerable work has been devoted to explaining predictive, deep learning-based models, and in turn how to evaluate explanations. An important class of evaluation methods are ones that are human-centere…

Experimental Designmodel

VākQA: A Benchmark and Evaluation Study for Telugu Spoken Factoid Question Answering

2026-09-17 · Bhavana Akkiraju, Ravi Sastry Kolluru, Sri Charan D, Srihari Bandarupalli 외 hf

Question answering has advanced rapidly with large language models, but predominantly for high-resource languages, in both text and spoken settings. Spoken question answering (SQA) benchmark for Telugu remains unexplored…

Question Answering