paper-with-me

Papers

FunBench: Benchmarking Fundus Reading Skills of MLLMs

2025-03-02 · Qijie Wei, Kaiheng Qian, Xirong Li

Multimodal Large Language Models (MLLMs) have shown significant potential in medical image analysis. However, their capabilities in interpreting fundus images, a critical skill for ophthalmology, remain under-evaluated. Existing benchmarks lack fine-grained task divisions and fail to provide modular analysis of its two key modules, i.e., large language model (LLM) and vision encoder (VE). This paper introduces FunBench, a novel visual question answering (VQA) benchmark designed to comprehensively evaluate MLLMs' fundus reading skills. FunBench features a hierarchical task organization across four levels (modality perception, anatomy perception, lesion analysis, and disease diagnosis). It also offers three targeted evaluation modes: linear-probe based VE evaluation, knowledge-prompted LLM evaluation, and holistic evaluation. Experiments on nine open-source MLLMs plus GPT-4o reveal significant deficiencies in fundus reading skills, particularly in basic tasks such as laterality recognition. The results highlight the limitations of current MLLMs and emphasize the need for domain-specific training and improved LLMs and VEs.

📄 PDF Abstract BibTeX arXiv:2503.00901

Code (0)

등록된 구현이 없습니다.

Tasks

AnatomyBenchmarkingLanguage ModelingLanguage ModellingLarge Language ModelMedical Image AnalysisQuestion AnsweringVisual Question AnsweringVisual Question Answering (VQA)

Similar Papers 제목 키워드 기반

Fundus-R1: Training a Fundus-Reading MLLM with Knowledge-Aware Reasoning on Public Data

2026-04-09 · Yuchuan Deng, Qijie Wei, Kaiheng Qian, Jiazhen Liu 외 arxiv

Fundus imaging such as CFP, OCT and UWF is crucial for the early detection of retinal anomalies and diseases. Fundus image understanding, due to its knowledge-intensive nature, poses a challenging vision-language task. A…

Reinforcement Learning

Constructing Ophthalmic MLLM for Positioning-diagnosis Collaboration Through Clinical Cognitive Chain Reasoning

2025-07-23 · Xinyao Liu, Diping Song arxiv

Multimodal large language models (MLLMs) demonstrate significant potential in the field of medical diagnosis. However, they face critical challenges in specialized domains such as ophthalmology, particularly the fragment…

Medical DiagnosisObject Detection

Assessing the Benchmarking Capacity of Machine Reading Comprehension Datasets

2019-11-21 · Saku Sugawara, Pontus Stenetorp, Kentaro Inui, Akiko Aizawa

Existing analysis work in machine reading comprehension (MRC) is largely concerned with evaluating the capabilities of systems. However, the capabilities of datasets are not assessed for benchmarking language understandi…

BenchmarkingMachine Reading ComprehensionReading ComprehensionSentence

Fantastic Questions and Where to Find Them: FairytaleQA – An Authentic Dataset for Narrative Comprehension

2022-05-01 · ACL 2022 5 · Ying Xu, Dakuo Wang, Mo Yu, Daniel Ritchie 외

Question answering (QA) is a fundamental means to facilitate assessment and training of narrative comprehension skills for both machines and young children, yet there is scarcity of high-quality QA datasets carefully des…

BenchmarkingQuestion AnsweringQuestion GenerationQuestion-Generation

Fantastic Questions and Where to Find Them: FairytaleQA -- An Authentic Dataset for Narrative Comprehension

2022-03-26 · Ying Xu, Dakuo Wang, Mo Yu, Daniel Ritchie 외

Question answering (QA) is a fundamental means to facilitate assessment and training of narrative comprehension skills for both machines and young children, yet there is scarcity of high-quality QA datasets carefully des…

BenchmarkingQuestion AnsweringQuestion GenerationQuestion-Generation