VisionUnite: A Vision-Language Foundation Model for Ophthalmology Enhanced with Clinical Knowledge
The need for improved diagnostic methods in ophthalmology is acute, especially in the less developed regions with limited access to specialists and advanced equipment. Therefore, we introduce VisionUnite, a novel vision-language foundation model for ophthalmology enhanced with clinical knowledge. VisionUnite has been pretrained on an extensive dataset comprising 1.24 million image-text pairs, and further refined using our proposed MMFundus dataset, which includes 296,379 high-quality fundus image-text pairs and 889,137 simulated doctor-patient dialogue instances. Our experiments indicate that VisionUnite outperforms existing generative foundation models such as GPT-4V and Gemini Pro. It also demonstrates diagnostic capabilities comparable to junior ophthalmologists. VisionUnite performs well in various clinical scenarios including open-ended multi-disease diagnosis, clinical explanation, and patient interaction, making it a highly versatile tool for initial ophthalmic disease screening. VisionUnite can also serve as an educational aid for junior ophthalmologists, accelerating their acquisition of knowledge regarding both common and rare ophthalmic conditions. VisionUnite represents a significant advancement in ophthalmology, with broad implications for diagnostics, medical education, and understanding of disease mechanisms.
Code (1)
Tasks
Clinical KnowledgeDiagnosticSimilar Papers 제목 키워드 기반
Language Enhanced Model for Eye (LEME): An Open-Source Ophthalmology-Specific Large Language Model
Large Language Models (LLMs) are poised to revolutionize healthcare. Ophthalmology-specific LLMs remain scarce and underexplored. We introduced an open-source, specialized LLM for ophthalmology, termed Language Enhanced …
AllLanguage ModelingLanguage ModellingLarge Language Model+2GlaBoost: A Multimodal Structured Framework for Glaucoma Risk Stratification
Early and accurate glaucoma detection is critical to prevent irreversible vision loss, yet existing AI methods often rely on unimodal inputs and lack interpretability. We present GlaBoost, a multimodal gradient boosting …
Feature ImportanceRET-CLIP: A Retinal Image Foundation Model Pre-trained with Clinical Diagnostic Reports
The Vision-Language Foundation model is increasingly investigated in the fields of computer vision and natural language processing, yet its exploration in ophthalmology and broader medical applications remains limited. T…
DiagnosticMulti-Label ClassificationMUlTI-LABEL-ClASSIFICATIONLMOD: A Large Multimodal Ophthalmology Dataset and Benchmark for Large Vision-Language Models
The prevalence of vision-threatening eye diseases is a significant global burden, with many cases remaining undiagnosed or diagnosed too late for effective treatment. Large vision-language models (LVLMs) have the potenti…
HallucinationA Survey of Multimodal Ophthalmic Diagnostics: From Task-Specific Approaches to Foundational Models
Visual impairment represents a major global health challenge, with multimodal imaging providing complementary information that is essential for accurate ophthalmic diagnosis. This comprehensive survey systematically revi…
Multimodal Deep LearningSelf-Supervised LearningReinforcement Learning