paper-with-me

홈 › Papers

Simulating Clinical AI Assistance using Multimodal LLMs: A Case Study in Diabetic Retinopathy

2025-09-16 · Nadim Barakat, William Lotter arxiv

Diabetic retinopathy (DR) is a leading cause of blindness worldwide, and AI systems can expand access to fundus photography screening. Current FDA-cleared systems primarily provide binary referral outputs, where this minimal output may limit clinical trust and utility. Yet, determining the most effective output format to enhance clinician-AI performance is an empirical challenge that is difficult to assess at scale. We evaluated multimodal large language models (MLLMs) for DR detection and their ability to simulate clinical AI assistance across different output types. Two models were tested on IDRiD and Messidor-2: GPT-4o, a general-purpose MLLM, and MedGemma, an open-source medical model. Experiments included: (1) baseline evaluation, (2) simulated AI assistance with synthetic predictions, and (3) actual AI-to-AI collaboration where GPT-4o incorporated MedGemma outputs. MedGemma outperformed GPT-4o at baseline, achieving higher sensitivity and AUROC, while GPT-4o showed near-perfect specificity but low sensitivity. Both models adjusted predictions based on simulated AI inputs, but GPT-4o's performance collapsed with incorrect ones, whereas MedGemma remained more stable. In actual collaboration, GPT-4o achieved strong results when guided by MedGemma's descriptive outputs, even without direct image access (AUROC up to 0.96). These findings suggest MLLMs may improve DR screening pipelines and serve as scalable simulators for studying clinical AI assistance across varying output configurations. Open, lightweight models such as MedGemma may be especially valuable in low-resource settings, while descriptive outputs could enhance explainability and clinician trust in clinical workflows.

📄 PDF Abstract BibTeX arXiv:2509.13234

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Are LLMs Ready to Assist Physicians? PhysAssistBench for Interactive Doctor-Patient-EHR Assistance

2026-06-17 · Tianming Du, Peijie Yu, Sihan Shang, Danli Shi 외 arxiv

The most plausible near-term role of medical LLMs is to assist rather than replace physicians, yet current evaluations often test isolated capabilities: clinical knowledge, EHR system interaction, or patient communicatio…

Clinical Knowledge

Can Multimodal Large Language Models Understand Pathologic Movements? A Pilot Study on Seizure Semiology

2026-05-05 · Lina Zhang, Tonmoy Monsoor, Mehmet Efe Lorasdagi, Prateik Sinha 외 arxiv

Multimodal Large Language Models (MLLMs) have demonstrated robust capabilities in recognizing everyday human activities, yet their potential for analyzing clinically significant involuntary movements in neurological diso…

Audio DenoisingPose Estimation

MTBBench: A Multimodal Sequential Clinical Decision-Making Benchmark in Oncology

2025-11-25 · Kiril Vasilev, Alexandre Misrahi, Eeshaan Jain, Phil F Cheng 외 arxiv

Multimodal Large Language Models (LLMs) hold promise for biomedical reasoning, but current benchmarks fail to capture the complexity of real-world clinical workflows. Existing evaluations primarily assess unimodal, decon…

ArogyaSutra: A Multi-Agent Framework for Multimodal Medical Reasoning in Indic Languages

2026-06-11 · Tanmoy Kanti Halder, Akash Ghosh, Subhadip Baidya, Arijit Roy 외 arxiv

Multimodal Large Language Models (MLLMs) have shown promising reasoning capabilities in general domains, yet their performance remains limited in specialized settings such as healthcare, especially in multilingual and lo…

Decision Making

EyePCR: A Comprehensive Benchmark for Fine-Grained Perception, Knowledge Comprehension and Clinical Reasoning in Ophthalmic Surgery

2025-09-19 · Gui Wang, Yang Wennuo, Xusen Ma, Zehao Zhong 외 arxiv

MLLMs (Multimodal Large Language Models) have showcased remarkable capabilities, but their performance in high-stakes, domain-specific scenarios like surgical settings, remains largely under-explored. To address this gap…

Clinical Knowledge