paper-with-me

Papers

Benchmarking Large Multimodal Models for Ophthalmic Visual Question Answering with OphthalWeChat

2025-05-26 · Pusheng Xu, Xia Gong, Xiaolan Chen, Weiyi Zhang, Jiancheng Yang, Bingjie Yan, Meng Yuan, Yalin Zheng, Mingguang He, Danli Shi

Purpose: To develop a bilingual multimodal visual question answering (VQA) benchmark for evaluating VLMs in ophthalmology. Methods: Ophthalmic image posts and associated captions published between January 1, 2016, and December 31, 2024, were collected from WeChat Official Accounts. Based on these captions, bilingual question-answer (QA) pairs in Chinese and English were generated using GPT-4o-mini. QA pairs were categorized into six subsets by question type and language: binary (Binary_CN, Binary_EN), single-choice (Single-choice_CN, Single-choice_EN), and open-ended (Open-ended_CN, Open-ended_EN). The benchmark was used to evaluate the performance of three VLMs: GPT-4o, Gemini 2.0 Flash, and Qwen2.5-VL-72B-Instruct. Results: The final OphthalWeChat dataset included 3,469 images and 30,120 QA pairs across 9 ophthalmic subspecialties, 548 conditions, 29 imaging modalities, and 68 modality combinations. Gemini 2.0 Flash achieved the highest overall accuracy (0.548), outperforming GPT-4o (0.522, P < 0.001) and Qwen2.5-VL-72B-Instruct (0.514, P < 0.001). It also led in both Chinese (0.546) and English subsets (0.550). Subset-specific performance showed Gemini 2.0 Flash excelled in Binary_CN (0.687), Single-choice_CN (0.666), and Single-choice_EN (0.646), while GPT-4o ranked highest in Binary_EN (0.717), Open-ended_CN (BLEU-1: 0.301; BERTScore: 0.382), and Open-ended_EN (BLEU-1: 0.183; BERTScore: 0.240). Conclusions: This study presents the first bilingual VQA benchmark for ophthalmology, distinguished by its real-world context and inclusion of multiple examinations per patient. The dataset reflects authentic clinical decision-making scenarios and enables quantitative evaluation of VLMs, supporting the development of accurate, specialized, and trustworthy AI systems for eye care.

📄 PDF Abstract BibTeX arXiv:2505.19624

Code (0)

등록된 구현이 없습니다.

Tasks

BenchmarkingQuestion AnsweringVisual Question AnsweringVisual Question Answering (VQA)

Similar Papers 제목 키워드 기반

Visual Question Answering in Ophthalmology: A Progressive and Practical Perspective

2024-10-22 · Xiaolan Chen, Ruoyu Chen, Pusheng Xu, Weiyi Zhang 외

Accurate diagnosis of ophthalmic diseases relies heavily on the interpretation of multimodal ophthalmic images, a process often time-consuming and expertise-dependent. Visual Question Answering (VQA) presents a potential…

Question AnsweringVisual Question AnsweringVisual Question Answering (VQA)

OphIn-500K: Curating Web-Scale Visual Instructions for Scaling Ophthalmic Multimodal Large Language Models

2026-05-27 · Xuanzhao Dong, Wenhui Zhu, Xiwen Chen, Hao Wang 외 arxiv

The advancement of general medical Multimodal Large Language Models (MLLMs) has shown great potential for building conversational assistants to support clinical diagnosis. However, their adaptation to highly specialized …

Visual Question Answering

OphGLM: Training an Ophthalmology Large Language-and-Vision Assistant based on Instructions and Dialogue

2023-06-21 · Weihao Gao, Zhuo Deng, Zhiyuan Niu, Fuju Rong 외

Large multimodal language models (LMMs) have achieved significant success in general domains. However, due to the significant differences between medical images and text and general web content, the performance of LMMs i…

Instruction FollowingLanguage ModelingLanguage ModellingLarge Language Model+1

EyeFound: A Multimodal Generalist Foundation Model for Ophthalmic Imaging

2024-05-18 · Danli Shi, Weiyi Zhang, Xiaolan Chen, Yexin Liu 외

Artificial intelligence (AI) is vital in ophthalmology, tackling tasks like diagnosis, classification, and visual question answering (VQA). However, existing AI models in this domain often require extensive annotation an…

Question AnsweringVisual Question AnsweringVisual Question Answering (VQA)

Towards Clinically Interpretable Ophthalmic VQA via Spatially-Grounded Lesion Evidence

2026-05-21 · Xingyue Wang, Bo Liu, Meng Wang, Zhixuan Zhang 외 arxiv

Visual Question Answering (VQA) holds great promise for clinical support, particularly in ophthalmology, where retinal fundus photography is essential for diagnosis. However, ophthalmic VQA benchmarks primarily emphasize…

Visual Question Answering