paper-with-me

홈 › Papers

EyeMVP: OCT-Informed Fundus Representation Learning via Paired CFP--OCT Pretraining

2026-06-13 · Zhuo Deng, Ruiheng Zhang, Ziheng Zhang, Weihao Gao, Yitong Li, Qian Wang, Lei Shao, Jiaoyue Dong, Zhixi Zeng, Lijian Fang, Haibo Wang, Xiaobin Lin, Tao Liu, Zhicheng Du, Zhengwei Zhang, Lin Yang, Zheng Gong, Xinyu Zhao, Zhenquan Wu, Fang Li, Zhiguang Zhou, Guoming Zhang, Sun Jing, Han Lv, Wenbin We, Lan Ma arxiv

Color fundus photography (CFP) is the mainstay of large-scale retinal screening, but its diagnostic capacity is limited by the lack of depth-resolved structure, which optical coherence tomography (OCT) provides yet is less accessible at population scale. We present EyeMVP, a cross-modal retinal foundation model that uses paired CFP--OCT pretraining to learn OCT-informed CFP representations while requiring only CFP at inference. Pretrained on 674,893 same-eye same-day CFP--OCT triples from 112,642 patients across eight hospitals, EyeMVP uses cross-modal masked reconstruction to enrich CFP features with OCT-associated supervision, and combines source-constrained cross-attention with CFP-derived structural masks to accommodate the non-aligned geometry of en-face CFP and cross-sectional OCT. Across 15 dataset-level settings spanning classification and segmentation, under both full-data and few-shot regimes, EyeMVP performs on par with or better than representative retinal foundation models, with consistent gains on macular and optic-nerve tasks; it attains AUROCs of 0.923 for macular edema and 0.867 for myopic macular schisis, two conditions poorly resolved in CFP. In an exploratory reader study, EyeMVP surpasses junior and intermediate ophthalmologists but not seniors on macular edema, while exceeding all groups on myopic macular schisis. These results indicate that cross-modal reconstruction can enrich CFP representations with OCT-associated supervision, offering a practical route to stronger CFP-based screening.

📄 PDF Abstract BibTeX arXiv:2606.15129

Code (0)

등록된 구현이 없습니다.

Tasks

Representation Learning

Similar Papers 제목 키워드 기반

Clinical Graph-Mediated Distillation for Unpaired MRI-to-CFI Hypertension Prediction

2026-03-23 · Dillan Imans, Phuoc-Nguyen Bui, Duc-Tai Le, Hyunseung Choo arxiv

Retinal fundus imaging enables low-cost and scalable hypertension (HTN) screening, but HTN-related retinal cues are subtle, yielding high-variance predictions. Brain MRI provides stronger vascular and small-vessel-diseas…

MM-Retinal V2: Transfer an Elite Knowledge Spark into Fundus Vision-Language Pretraining

2025-01-27 · Ruiqi Wu, Na Su, Chenran Zhang, Tengfei Ma 외

Vision-language pretraining (VLP) has been investigated to generalize across diverse downstream tasks for fundus image analysis. Although recent methods showcase promising achievements, they significantly rely on large-s…

Contrastive LearningTransfer Learning

Towards Interpretable Foundation Models for Retinal Fundus Images

2026-03-19 · Samuel Ofosu Mensah, Camila Roa, Kerol Djoumessi, Philipp Berens arxiv

Foundation models are used to extract transferable representations from large amounts of unlabeled data, typically via self-supervised learning (SSL). However, many of these models rely on architectures that offer limite…

Self-Supervised Learning

EyeBench: A Call for More Rigorous Evaluation of Retinal Image Enhancement

2025-02-20 · Wenhui Zhu, Xuanzhao Dong, Xin Li, Yujian Xiong 외

Over the past decade, generative models have achieved significant success in enhancement fundus images.However, the evaluation of these models still presents a considerable challenge. A comprehensive evaluation benchmark…

DenoisingImage EnhancementLesion SegmentationSSIM

Enhancing Diagnostic Accuracy in Rare and Common Fundus Diseases with a Knowledge-Rich Vision-Language Model

2024-06-13 · Meng Wang, Tian Lin, Aidi Lin, Kai Yu 외

Previous foundation models for fundus images were pre-trained with limited disease categories and knowledge base. Here we introduce a knowledge-rich vision-language model (RetiZero) that leverages knowledge from more tha…

DiagnosticImage RetrievalLanguage ModelingLanguage Modelling+1