paper-with-me

홈 › Papers

MedBookVQA: A Systematic and Comprehensive Medical Benchmark Derived from Open-Access Book

2025-06-01 · Sau Lai Yip, Sunan He, Yuxiang Nie, Shu Pui Chan, Yilin Ye, Sum Ying Lam, Hao Chen

The accelerating development of general medical artificial intelligence (GMAI), powered by multimodal large language models (MLLMs), offers transformative potential for addressing persistent healthcare challenges, including workforce deficits and escalating costs. The parallel development of systematic evaluation benchmarks emerges as a critical imperative to enable performance assessment and provide technological guidance. Meanwhile, as an invaluable knowledge source, the potential of medical textbooks for benchmark development remains underexploited. Here, we present MedBookVQA, a systematic and comprehensive multimodal benchmark derived from open-access medical textbooks. To curate this benchmark, we propose a standardized pipeline for automated extraction of medical figures while contextually aligning them with corresponding medical narratives. Based on this curated data, we generate 5,000 clinically relevant questions spanning modality recognition, disease classification, anatomical identification, symptom diagnosis, and surgical procedures. A multi-tier annotation system categorizes queries through hierarchical taxonomies encompassing medical imaging modalities (42 categories), body anatomies (125 structures), and clinical specialties (31 departments), enabling nuanced analysis across medical subdomains. We evaluate a wide array of MLLMs, including proprietary, open-sourced, medical, and reasoning models, revealing significant performance disparities across task types and model categories. Our findings highlight critical capability gaps in current GMAI systems while establishing textbook-derived multimodal benchmarks as essential evaluation tools. MedBookVQA establishes textbook-derived benchmarking as a critical paradigm for advancing clinical AI, exposing limitations in GMAI systems while providing anatomically structured performance metrics across specialties.

📄 PDF Abstract BibTeX arXiv:2506.00855

Code (1)

slyipae1/MedBookVQA 공식 구현 pytorch

Tasks

Benchmarking

Similar Papers 제목 키워드 기반

MedHallu: A Comprehensive Benchmark for Detecting Medical Hallucinations in Large Language Models

2025-02-20 · Shrey Pandit, Jiawei Xu, Junyuan Hong, Zhangyang Wang 외

Advancements in Large Language Models (LLMs) and their increasing use in medical question-answering necessitate rigorous evaluation of their reliability. A critical challenge lies in hallucination, where models generate …

Decision MakingHallucinationMedical Question AnsweringQuestion Answering

Medical Reasoning with Large Language Models: A Survey and MR-Bench

2026-03-17 · Xiaohan Ren, Chenxiao Fan, Wenyin Ma, Hongliang He 외 arxiv

Large language models (LLMs) have achieved strong performance on medical exam-style tasks, motivating growing interest in their deployment in real-world clinical settings. However, clinical decision-making is inherently …

BRIDGE: Benchmarking Large Language Models for Understanding Real-world Clinical Practice Text

2025-04-28 · Jiageng Wu, Bowen Gu, Ren Zhou, Kevin Xie 외

Large language models (LLMs) hold great promise for medical applications and are evolving rapidly, with new models being released at an accelerated pace. However, current evaluations of LLMs in clinical contexts remain l…

Benchmarking

MedUAG: Unified Understanding and Generation for Medical Multimodal Models

2026-08-19 · Zijie Meng, Yuncheng Zhang, Hualiang Wang, Yitian Tang 외 arxiv

Recent Multimodal Large Language Models (MLLMs) are rapidly evolving into unified understanding and generation (UAG) frameworks. However, extending these unified paradigms to the medical domain is hindered by: the absenc…

Benchmark Evaluation of Federated Learning on Multi-organ Images

2026-07-09 · Junbin Mao, Xu Tian, Jianchun Zhu, Ludi Li 외 arxiv

The privacy requirements of medical data and its substantial variations across organs and modalities hinder the clinical implementation of medical AI. Federated learning (FL) is a feasible approach to overcome these chal…

Federated Learning