paper-with-me

Papers

VisionUnite: A Vision-Language Foundation Model for Ophthalmology Enhanced with Clinical Knowledge

2024-08-05 · Zihan Li, Diping Song, Zefeng Yang, Deming Wang, Fei Li, Xiulan Zhang, Paul E. Kinahan, Yu Qiao

The need for improved diagnostic methods in ophthalmology is acute, especially in the less developed regions with limited access to specialists and advanced equipment. Therefore, we introduce VisionUnite, a novel vision-language foundation model for ophthalmology enhanced with clinical knowledge. VisionUnite has been pretrained on an extensive dataset comprising 1.24 million image-text pairs, and further refined using our proposed MMFundus dataset, which includes 296,379 high-quality fundus image-text pairs and 889,137 simulated doctor-patient dialogue instances. Our experiments indicate that VisionUnite outperforms existing generative foundation models such as GPT-4V and Gemini Pro. It also demonstrates diagnostic capabilities comparable to junior ophthalmologists. VisionUnite performs well in various clinical scenarios including open-ended multi-disease diagnosis, clinical explanation, and patient interaction, making it a highly versatile tool for initial ophthalmic disease screening. VisionUnite can also serve as an educational aid for junior ophthalmologists, accelerating their acquisition of knowledge regarding both common and rare ophthalmic conditions. VisionUnite represents a significant advancement in ophthalmology, with broad implications for diagnostics, medical education, and understanding of disease mechanisms.

📄 PDF Abstract BibTeX arXiv:2408.02865

Code (1)

HUANGLIZI/VisionUnite 공식 구현 pytorch

Tasks

Clinical KnowledgeDiagnostic

Similar Papers 제목 키워드 기반

Language Enhanced Model for Eye (LEME): An Open-Source Ophthalmology-Specific Large Language Model

2024-10-01 · Aidan Gilson, Xuguang Ai, Qianqian Xie, Sahana Srinivasan 외

Large Language Models (LLMs) are poised to revolutionize healthcare. Ophthalmology-specific LLMs remain scarce and underexplored. We introduced an open-source, specialized LLM for ophthalmology, termed Language Enhanced …

AllLanguage ModelingLanguage ModellingLarge Language Model+2

GlaBoost: A Multimodal Structured Framework for Glaucoma Risk Stratification

2025-08-03 · Cheng Huang, Zeyu Han, Weizheng Xie, Karanjit Kooner 외 arxiv

Early and accurate glaucoma detection is critical to prevent irreversible vision loss, yet existing AI methods often rely on unimodal inputs and lack interpretability. We present GlaBoost, a multimodal gradient boosting …

Feature Importance

RET-CLIP: A Retinal Image Foundation Model Pre-trained with Clinical Diagnostic Reports

2024-05-23 · Jiawei Du, Jia Guo, Weihang Zhang, Shengzhu Yang 외

The Vision-Language Foundation model is increasingly investigated in the fields of computer vision and natural language processing, yet its exploration in ophthalmology and broader medical applications remains limited. T…

DiagnosticMulti-Label ClassificationMUlTI-LABEL-ClASSIFICATION

LMOD: A Large Multimodal Ophthalmology Dataset and Benchmark for Large Vision-Language Models

2024-10-02 · Zhenyue Qin, Yu Yin, Dylan Campbell, Xuansheng Wu 외

The prevalence of vision-threatening eye diseases is a significant global burden, with many cases remaining undiagnosed or diagnosed too late for effective treatment. Large vision-language models (LVLMs) have the potenti…

Hallucination

A Survey of Multimodal Ophthalmic Diagnostics: From Task-Specific Approaches to Foundational Models

2025-07-31 · Xiaoling Luo, Ruli Zheng, Qiaojian Zheng, Zibo Du 외 arxiv

Visual impairment represents a major global health challenge, with multimodal imaging providing complementary information that is essential for accurate ophthalmic diagnosis. This comprehensive survey systematically revi…

Multimodal Deep LearningSelf-Supervised LearningReinforcement Learning