paper-with-me

Papers

A Study on Multimodal and Interactive Explanations for Visual Question Answering

2020-03-01 · Kamran Alipour, Jurgen P. Schulze, Yi Yao, Avi Ziskind, Giedrius Burachas

Explainability and interpretability of AI models is an essential factor affecting the safety of AI. While various explainable AI (XAI) approaches aim at mitigating the lack of transparency in deep networks, the evidence of the effectiveness of these approaches in improving usability, trust, and understanding of AI systems are still missing. We evaluate multimodal explanations in the setting of a Visual Question Answering (VQA) task, by asking users to predict the response accuracy of a VQA agent with and without explanations. We use between-subjects and within-subjects experiments to probe explanation effectiveness in terms of improving user prediction accuracy, confidence, and reliance, among other factors. The results indicate that the explanations help improve human prediction accuracy, especially in trials when the VQA system's answer is inaccurate. Furthermore, we introduce active attention, a novel method for evaluating causal attentional effects through intervention by editing attention maps. User explanation ratings are strongly correlated with human prediction accuracy and suggest the efficacy of these explanations in human-machine AI collaboration tasks.

📄 PDF Abstract BibTeX arXiv:2003.00431

Code (0)

등록된 구현이 없습니다.

Tasks

Explainable Artificial Intelligence (XAI)PredictionQuestion AnsweringVisual Question AnsweringVisual Question Answering (VQA)

Methods 이 논문이 사용한 방법론

Interpretability 설명 없음

Similar Papers 제목 키워드 기반

LININ: Logic Integrated Neural Inference Network for Explanatory Visual Question Answering

2024-12-24 · IEEE Transactions on Multimedia 2024 12 · Dizhan Xue, Shengsheng Qian, Quan Fang, Changsheng Xu

Explanatory Visual Question Answering (EVQA) is a recently proposed multimodal reasoning task consisting of answering the visual question and generating multimodal explanations for the reasoning processes. Unlike traditi…

Explanatory Visual Question AnsweringMultimodal ReasoningQuestion AnsweringVisual Question Answering+1

Few-Shot Multimodal Explanation for Visual Question Answering

2024-10-28 · ACM MM 2024 10 · Dizhan Xue, Shengsheng Qian, Changsheng Xu

A key object in eXplainable Artificial Intelligence (XAI) is to create intelligent systems capable of reasoning and explaining real-world data to facilitate reliable decision-making. Recent studies have acknowledged the …

Explainable artificial intelligenceExplainable Artificial Intelligence (XAI)FS-MEVQAQuestion Answering+3

Variational Causal Inference Network for Explanatory Visual Question Answering

2023-01-01 · ICCV 2023 1 · Dizhan Xue, Shengsheng Qian, Changsheng Xu

Explanatory Visual Question Answering (EVQA) is a recently proposed multimodal reasoning task that requires answering visual questions and generating multimodal explanations for the reasoning processes. Unlike tradit…

Explanation GenerationExplanatory Visual Question AnsweringFS-MEVQAMultimodal Reasoning+3

Interactive Sketchpad: A Multimodal Tutoring System for Collaborative, Visual Problem-Solving

2025-02-12 · Steven-Shine Chen, JiMin Lee, Paul Pu Liang

Humans have long relied on visual aids like sketches and diagrams to support reasoning and problem-solving. Visual tools, like auxiliary lines in geometry or graphs in calculus, are essential for understanding complex id…

Mathmultimodal interaction

Mobile App Tasks with Iterative Feedback (MoTIF): Addressing Task Feasibility in Interactive Visual Environments

2021-04-17 · Andrea Burns, Deniz Arsan, Sanjna Agrawal, Ranjitha Kumar 외

In recent years, vision-language research has shifted to study tasks which require more complex reasoning, such as interactive question answering, visual common sense reasoning, and question-answer plausibility predictio…

Common Sense ReasoningQuestion Answering