paper-with-me

Papers

Are Multimodal LLMs Ready for Clinical Dermatology? A Real-World Evaluation in Dermatology

2026-05-01 · Roy Jiang, Hyunjae Kim, Zhenyue Qin, Morten Lee, Margaret MacGibeny, Ailish Hanly, Angela Sadlowski, Shanin Chowdhury, Xuguang Ai, Jeffrey Gehlhausen, Qingyu Chen arxiv

Multimodal large language models (MLLMs) have demonstrated promise on publicly available dermatology benchmarks. However, benchmark performance may not generalize to real-world dermatologic decision-making. To quantify this benchmark-to-bedside gap, we evaluated four open-weight MLLMs (InternVL-Chat v1.5, LLaVA-Med v1.5, SkinGPT4 and MedGemma-4B-Instruct) and one commercial MLLM (GPT-4.1) across three publicly available dermatology datasets and a retrospective multi-site hospital-based dermatology consultation cohort comprising 5,811 cases and 46,405 clinical images. Models were evaluated on two clinically relevant tasks: differential diagnosis generation and severity-based triage. Diagnostic performance was modest on public datasets and declined substantially in the real-world cohort. On public benchmarks, top-3 diagnostic accuracy reached 26.55% for the best open-weight model and 42.25% for GPT-4.1. On real-world consultation cases using images alone, top-3 diagnostic accuracy fell to 1.50%-13.35% among open-weight models and 24.65% for GPT-4.1. Incorporating clinical context improved performance across all models, increasing top-3 diagnostic accuracy up to 28.75% among open-weight models and 38.93% for GPT-4.1. However, model outputs were highly sensitive to incomplete or erroneous consultation context. For severity-based triage, models achieved moderate sensitivity (above 60%), suggesting potential utility for screening but insufficient reliability for clinical deployment. These findings demonstrate that benchmark performance substantially overestimates the real-world clinical capability of current dermatology MLLMs.

📄 PDF Abstract BibTeX arXiv:2605.04098

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Towards Trustworthy Dermatology MLLMs: A Benchmark and Multimodal Evaluator for Diagnostic Narratives

2025-11-12 · Yuhao Shen, Jiahe Qian, Shuping Zhang, Zhangtianyi Chen 외 arxiv

Multimodal large language models (LLMs) are increasingly used to generate dermatology diagnostic narratives directly from images. However, reliable evaluation remains the primary bottleneck for responsible clinical deplo…

MM-Skin: Enhancing Dermatology Vision-Language Model with an Image-Text Dataset Derived from Textbooks

2025-05-09 · Wenqi Zeng, Yuqi Sun, Chenxi Ma, Weimin Tan 외

Medical vision-language models (VLMs) have shown promise as clinical assistants across various medical fields. However, specialized dermatology VLM capable of delivering professional and detailed diagnostic analysis rema…

DiagnosticInstruction FollowingLanguage ModelingLanguage Modelling+4

Evaluating Strategies for Synthesizing Clinical Notes for Medical Multimodal AI

2025-11-26 · Niccolo Marini, Zhaohui Liang, Sivaramakrishnan Rajaraman, Zhiyun Xue 외 arxiv

Multimodal (MM) learning is emerging as a promising paradigm in biomedical artificial intelligence (AI) applications, integrating complementary modality, which highlight different aspects of patient health. The scarcity …

Cross-Modal Retrieval

A Multimodal Vision Foundation Model for Clinical Dermatology

2024-10-19 · Siyuan Yan, Zhen Yu, Clare Primiero, Cristina Vico-Alonso 외

Diagnosing and treating skin diseases require advanced visual skills across domains and the ability to synthesize information from multiple imaging modalities. While current deep learning models excel at specific tasks l…

DiagnosticLesion SegmentationmodelPrognosis+2

Towards Realization of Augmented Intelligence in Dermatology: Advances and Future Directions

2021-05-21 · Roxana Daneshjou, Carrie Kovarik, Justin M Ko

Artificial intelligence (AI) algorithms using deep learning have advanced the classification of skin disease images; however these algorithms have been mostly applied "in silico" and not validated clinically. Most dermat…

Binary ClassificationDeep LearningDiagnosticPosition