paper-with-me

홈 › Papers

A Systematic Evaluation of GPT-4V's Multimodal Capability for Medical Image Analysis

2023-10-31 · Yingshu Li, Yunyi Liu, Zhanyu Wang, Xinyu Liang, Lei Wang, Lingqiao Liu, Leyang Cui, Zhaopeng Tu, Longyue Wang, Luping Zhou

This work conducts an evaluation of GPT-4V's multimodal capability for medical image analysis, with a focus on three representative tasks of radiology report generation, medical visual question answering, and medical visual grounding. For the evaluation, a set of prompts is designed for each task to induce the corresponding capability of GPT-4V to produce sufficiently good outputs. Three evaluation ways including quantitative analysis, human evaluation, and case study are employed to achieve an in-depth and extensive evaluation. Our evaluation shows that GPT-4V excels in understanding medical images and is able to generate high-quality radiology reports and effectively answer questions about medical images. Meanwhile, it is found that its performance for medical visual grounding needs to be substantially improved. In addition, we observe the discrepancy between the evaluation outcome from quantitative analysis and that from human evaluation. This discrepancy suggests the limitations of conventional metrics in assessing the performance of large language models like GPT-4V and the necessity of developing new metrics for automatic quantitative analysis.

📄 PDF Abstract BibTeX arXiv:2310.20381

Code (0)

등록된 구현이 없습니다.

Tasks

DescriptiveMedical Image AnalysisMedical Visual Question AnsweringQuestion AnsweringVisual GroundingVisual Question AnsweringVisual Question Answering (VQA)

Methods 이 논문이 사용한 방법론

SET Dynamic Sparse Training method where weight mask is updated randomly periodically
Focus 설명 없음

Similar Papers 제목 키워드 기반

MMRareBench: A Rare-Disease Multimodal and Multi-Image Medical Benchmark

2026-04-12 · Junzhi Ning, Jiashi Lin, Yingying Fang, Wei Li 외 arxiv

Multimodal large language models (MLLMs) have advanced clinical tasks for common conditions, but their performance on rare diseases remains largely untested. In rare-disease scenarios, clinicians often lack prior clinica…

Clinical Knowledge

MedQ-Deg: A Multidimensional Benchmark for Evaluating MLLMs Across Medical Image Quality Degradations

2026-03-08 · Jiyao Liu, Junzhi Ning, Chenglong Ma, Wanying Qu 외 arxiv

Despite impressive performance on standard benchmarks, multimodal large language models (MLLMs) face critical challenges in real-world clinical environments where medical images inevitably suffer various quality degradat…

A multimodal slice discovery framework for systematic failure detection and explanation in medical image classification

2026-02-27 · Yixuan Liu, Kanwal K. Bhatia, Ahmed E. Fetit arxiv

Despite advances in machine learning-based medical image classifiers, the safety and reliability of these systems remain major concerns in practical settings. Existing auditing approaches mainly rely on unimodal features…

Medical Image ClassificationExplanation Generation

MedBookVQA: A Systematic and Comprehensive Medical Benchmark Derived from Open-Access Book

2025-06-01 · Sau Lai Yip, Sunan He, Yuxiang Nie, Shu Pui Chan 외

The accelerating development of general medical artificial intelligence (GMAI), powered by multimodal large language models (MLLMs), offers transformative potential for addressing persistent healthcare challenges, includ…

Benchmarking

Can Multimodal Large Language Models Understand OCT?

2026-07-18 · Baochen Fu, Wenzhi Deng, Baihao Jin, Yang Li 외 hf

Optical coherence tomography (OCT) imaging is essential for the diagnosis and treatment of retinal diseases. Although multimodal large language models (MLLMs) have demonstrated considerable potential in medical image ana…

Visual Question AnsweringDomain Adaptation