paper-with-me

홈 › Papers

MedThink: Explaining Medical Visual Question Answering via Multimodal Decision-Making Rationale

2024-04-18 · Xiaotang Gai, Chenyi Zhou, Jiaxiang Liu, Yang Feng, Jian Wu, Zuozhu Liu

Medical Visual Question Answering (MedVQA), which offers language responses to image-based medical inquiries, represents a challenging task and significant advancement in healthcare. It assists medical experts to swiftly interpret medical images, thereby enabling faster and more accurate diagnoses. However, the model interpretability and transparency of existing MedVQA solutions are often limited, posing challenges in understanding their decision-making processes. To address this issue, we devise a semi-automated annotation process to streamline data preparation and build new benchmark MedVQA datasets R-RAD, R-SLAKE and R-Path. These datasets provide intermediate medical decision-making rationales generated by multimodal large language models and human annotations for question-answering pairs in existing MedVQA datasets, i.e., VQA-RAD, SLAKE and PathVQA. Moreover, we design a novel framework, MedThink, which finetunes lightweight pretrained generative models by incorporating medical decision-making rationales. MedThink includes three distinct strategies to generate decision outcomes and corresponding rationales, thereby clearly showcasing the medical decision-making process during reasoning. Our comprehensive experiments show that our method achieves an accuracy of 83.5% on R-RAD, 86.3% on R-SLAKE and 87.2% on R-Path. These results significantly exceed those of existing state-of-the-art models with comparable parameters. Datasets and code will be released.

📄 PDF Abstract BibTeX arXiv:2404.12372

Code (0)

등록된 구현이 없습니다.

Tasks

Decision MakingMedical Visual Question AnsweringQuestion AnsweringVisual Question AnsweringVisual Question Answering (VQA)

Similar Papers 제목 키워드 기반

Automating Expert-Level Medical Reasoning Evaluation of Large Language Models

2025-07-10 · Shuang Zhou, Wenya Xie, Jiaxi Li, Zaifu Zhan 외 arxiv

As large language models (LLMs) become increasingly integrated into clinical decision-making, ensuring transparent and trustworthy reasoning is essential. However, existing evaluation strategies of LLMs' medical reasonin…

Vision-Language Models for Medical Report Generation and Visual Question Answering: A Review

2024-03-04 · Iryna Hartsock, Ghulam Rasool

Medical vision-language models (VLMs) combine computer vision (CV) and natural language processing (NLP) to analyze visual and textual medical data. Our paper reviews recent advancements in developing VLMs specialized fo…

Medical Report GenerationQuestion AnsweringVisual Question AnsweringVisual Question Answering (VQA)

MedThink: Enhancing Diagnostic Accuracy in Small Models via Teacher-Guided Reasoning Correction

2026-04-09 · Xinchun Su, Chunxu Luo, Lipeng Ma, Yixuan Li 외 arxiv

Accurate clinical diagnosis requires extensive domain knowledge and complex clinical reasoning capabilities. Although large language models (LLMs) hold great potential for clinical reasoning, their high computational and…

Computational EfficiencyKnowledge Distillation

Benchmarking Large Language Models on Answering and Explaining Challenging Medical Questions

2024-02-28 · Hanjie Chen, Zhouxiang Fang, Yash Singla, Mark Dredze

LLMs have demonstrated impressive performance in answering medical questions, such as achieving passing scores on medical licensing examinations. However, medical board exams or general clinical questions do not capture …

BenchmarkingMultiple-choiceQuestion Answering

Structure Causal Models and LLMs Integration in Medical Visual Question Answering

2025-05-05 · Zibo Xu, Qiang Li, Weizhi Nie, Weijie Wang 외

Medical Visual Question Answering (MedVQA) aims to answer medical questions according to medical images. However, the complexity of medical data leads to confounders that are difficult to observe, so bias between images …

Causal InferenceMedical Visual Question AnsweringQuestion AnsweringVisual Question Answering