paper-with-me

홈 › Papers

MedXplain-VQA: Multi-Component Explainable Medical Visual Question Answering

2025-10-26 · Hai-Dang Nguyen, Minh-Anh Dang, Minh-Tan Le, Minh-Tuan Le arxiv

Explainability is critical for the clinical adoption of medical visual question answering (VQA) systems, as physicians require transparent reasoning to trust AI-generated diagnoses. We present MedXplain-VQA, a comprehensive framework integrating five explainable AI components to deliver interpretable medical image analysis. The framework leverages a fine-tuned BLIP-2 backbone, medical query reformulation, enhanced Grad-CAM attention, precise region extraction, and structured chain-of-thought reasoning via multi-modal language models. To evaluate the system, we introduce a medical-domain-specific framework replacing traditional NLP metrics with clinically relevant assessments, including terminology coverage, clinical structure quality, and attention region relevance. Experiments on 500 PathVQA histopathology samples demonstrate substantial improvements, with the enhanced system achieving a composite score of 0.683 compared to 0.378 for baseline methods, while maintaining high reasoning confidence (0.890). Our system identifies 3-5 diagnostically relevant regions per sample and generates structured explanations averaging 57 words with appropriate clinical terminology. Ablation studies reveal that query reformulation provides the most significant initial improvement, while chain-of-thought reasoning enables systematic diagnostic processes. These findings underscore the potential of MedXplain-VQA as a robust, explainable medical VQA system. Future work will focus on validation with medical experts and large-scale clinical datasets to ensure clinical readiness.

📄 PDF Abstract BibTeX arXiv:2510.22803

Code (0)

등록된 구현이 없습니다.

Tasks

Visual Question Answering

Similar Papers 제목 키워드 기반

GEMeX: A Large-Scale, Groundable, and Explainable Medical VQA Benchmark for Chest X-ray Diagnosis

2024-11-25 · Bo Liu, Ke Zou, LiMing Zhan, Zexin Lu 외

Medical Visual Question Answering (VQA) is an essential technology that integrates computer vision and natural language processing to automatically respond to clinical inquiries about medical images. However, current med…

Medical Visual Question AnsweringMultiple-choiceQuestion AnsweringVisual Question Answering+1

VALD-MD: Visual Attribution via Latent Diffusion for Medical Diagnostics

2024-01-02 · Ammar A. Siddiqui, Santosh Tirunagari, Tehseen Zia, David Windridge

Visual attribution in medical imaging seeks to make evident the diagnostically-relevant components of a medical image, in contrast to the more common detection of diseased tissue deployed in standard machine vision pipel…

MS-SSIMSSIM

Lesion Guided Explainable Few Weak-shot Medical Report Generation

2022-11-16 · Jinghan Sun, Dong Wei, Liansheng Wang, Yefeng Zheng

Medical images are widely used in clinical practice for diagnosis. Automatically generating interpretable medical reports can reduce radiologists' burden and facilitate timely care. However, most existing approaches to a…

Medical Report Generation

Developing A Visual-Interactive Interface for Electronic Health Record Labeling: An Explainable Machine Learning Approach

2022-09-26 · Donlapark Ponnoprat, Parichart Pattarapanitchai, Phimphaka Taninpong, Suthep Suantai 외

Labeling a large number of electronic health records is expensive and time consuming, and having a labeling assistant tool can significantly reduce medical experts' workload. Nevertheless, to gain the experts' trust, the…

MICA: Towards Explainable Skin Lesion Diagnosis via Multi-Level Image-Concept Alignment

2024-01-16 · Yequan Bie, Luyang Luo, Hao Chen

Black-box deep learning approaches have showcased significant potential in the realm of medical image analysis. However, the stringent trustworthiness requirements intrinsic to the medical field have catalyzed research i…

Concept AlignmentExplainable artificial intelligenceExplainable Artificial Intelligence (XAI)Medical Image Analysis