paper-with-me

홈 › Papers

MedKCO: Medical Vision-Language Pretraining via Knowledge-Driven Cognitive Orchestration

2026-03-10 · Chenran Zhang, Ruiqi Wu, Tao Zhou, Yi Zhou arxiv

Medical vision-language pretraining (VLP) models have recently been investigated for their generalization to diverse downstream tasks. However, current medical VLP methods typically force the model to learn simple and complex concepts simultaneously. This anti-cognitive process leads to suboptimal feature representations, especially under distribution shift. To address this limitation, we propose a Knowledge-driven Cognitive Orchestration for Medical VLP (MedKCO) that involves both the ordering of the pretraining data and the learning objective of vision-language contrast. Specifically, we design a two level curriculum by incorporating diagnostic sensitivity and intra-class sample representativeness for the ordering of the pretraining data. Moreover, considering the inter-class similarity of medical images, we introduce a self-paced asymmetric contrastive loss to dynamically adjust the participation of the pretraining objective. We evaluate the proposed pretraining method on three medical imaging scenarios in multiple vision-language downstream tasks, and compare it with several curriculum learning methods. Extensive experiments show that our method significantly surpasses all baselines. https://github.com/Mr-Talon/MedKCO.

📄 PDF Abstract BibTeX arXiv:2603.09101

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Medical Vision Language Pretraining: A survey

2023-12-11 · Prashant Shrestha, Sanskar Amgain, Bidur Khanal, Cristian A. Linte 외

Medical Vision Language Pretraining (VLP) has recently emerged as a promising solution to the scarcity of labeled data in the medical domain. By leveraging paired/unpaired vision and text datasets through self-supervised…

Self-Supervised LearningSurvey

PhenoLIP: Integrating Phenotype Ontology Knowledge into Medical Vision-Language Pretraining

2026-02-05 · Cheng Liang, Chaoyi Wu, Weike Zhao, Ya Zhang 외 arxiv

Recent progress in large-scale CLIP-like vision-language models(VLMs) has greatly advanced medical image analysis. However, most existing medical VLMs still rely on coarse image-text contrastive objectives and fail to ca…

Phenotype classificationKnowledge DistillationCross-Modal Retrieval

MedTri: A Platform for Structured Medical Report Normalization to Enhance Vision-Language Pretraining

2026-02-25 · Yuetan Chu, Xinhua Ma, Xinran Jin, Gongning Luo 외 arxiv

Medical vision-language pretraining increasingly relies on medical reports as large-scale supervisory signals; however, raw reports often exhibit substantial stylistic heterogeneity, variable length, and a considerable a…

Multi-Aspect Knowledge-Enhanced Medical Vision-Language Pretraining with Multi-Agent Data Generation

2025-12-03 · Xieji Li, Siyuan Yan, Yingsheng Liu, H. Peter Soyer 외 arxiv

Vision-language pretraining (VLP) has emerged as a powerful paradigm in medical image analysis, enabling representation learning from large-scale image-text pairs without relying on expensive manual annotations. However,…

Representation LearningCross-Modal Retrieval

Medical Vision-Language Pre-Training for Brain Abnormalities

2024-04-27 · Masoud Monajatipoor, Zi-Yi Dou, Aichi Chien, Nanyun Peng 외

Vision-language models have become increasingly powerful for tasks that require an understanding of both visual and linguistic elements, bridging the gap between these modalities. In the context of multimodal clinical AI…

Language ModelingLanguage Modelling