paper-with-me

Papers

Multi-modal Vision Pre-training for Medical Image Analysis

2025-01-01 · CVPR 2025 1 · Shaohao Rui, Lingzhi Chen, Zhenyu Tang, Lilong Wang, Mianxin Liu, Shaoting Zhang, Xiaosong Wang

Self-supervised learning has greatly facilitated medical image analysis by suppressing the training data requirement for real-world applications. Current paradigms predominantly rely on self-supervision within uni-modal image data, thereby neglecting the inter-modal correlations essential for effective learning of cross-modal image representations. This limitation is particularly significant for naturally grouped multi-modal data, e.g., multi-parametric MRI scans for a patient undergoing various functional imaging protocols in the same study. To bridge this gap, we conduct a novel multi-modal image pre-training with three proxy tasks to facilitate the learning of cross-modality representations and correlations using multi-modal brain MRI scans (over 2.4 million images in 16,022 scans of 3,755 patients), i.e., cross-modal image reconstruction, modality-aware contrastive learning, and modality template distillation. To demonstrate the generalizability of our pre-trained model, we conduct extensive experiments on various benchmarks with ten downstream tasks. The superior performance of our method is reported in comparison to state-of-the-art pre-training methods, with Dice Score improvement of 0.28%-14.47% across six segmentation benchmarks and a consistent accuracy boost of 0.65%-18.07% in four individual image classification tasks.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Tasks

Contrastive Learningimage-classificationImage ClassificationImage ReconstructionMedical Image AnalysisSelf-Supervised Learning

Similar Papers 제목 키워드 기반

Uni-Mlip: Unified Self-supervision for Medical Vision Language Pre-training

2024-11-20 · Ameera Bawazir, Kebin Wu, Wenbin Li

Recent advancements in vision-language pre-training via contrastive learning have significantly improved performance across computer vision tasks. However, in the medical domain, obtaining multimodal data is often costly…

Contrastive Learningimage-classificationImage ClassificationImage-text Retrieval+5

Multi-modal Understanding and Generation for Medical Images and Text via Vision-Language Pre-Training

2021-05-24 · Jong Hak Moon, Hyungyung Lee, Woncheol Shin, Young-Hak Kim 외

Recently a number of studies demonstrated impressive performance on diverse vision-language multi-modal tasks such as image captioning and visual question answering by extending the BERT architecture with multi-modal pre…

Image CaptioningMedical Visual Question AnsweringMultimodal Deep LearningQuestion Answering+5

MISS: A Generative Pretraining and Finetuning Approach for Med-VQA

2024-01-10 · Jiawei Chen, Dingkang Yang, Yue Jiang, Yuxuan Lei 외

Medical visual question answering (VQA) is a challenging multimodal task, where Vision-Language Pre-training (VLP) models can effectively improve the generalization performance. However, most methods in the medical field…

Medical Visual Question AnsweringMulti-Task LearningQuestion AnsweringSelf-Supervised Learning+2

GMAI-VL & GMAI-VL-5.5M: A Large Vision-Language Model and A Comprehensive Multimodal Dataset Towards General Medical AI

2024-11-21 · Tianbin Li, Yanzhou Su, Wei Li, Bin Fu 외

Despite significant advancements in general AI, its effectiveness in the medical domain is limited by the lack of specialized medical knowledge. To address this, we formulate GMAI-VL-5.5M, a multimodal medical dataset cr…

Decision MakingLanguage ModelingLanguage ModellingQuestion Answering+1

ZALM3: Zero-Shot Enhancement of Vision-Language Alignment via In-Context Information in Multi-Turn Multimodal Medical Dialogue

2024-09-26 · Zhangpu Li, Changhong Zou, Suxue Ma, Zhicheng Yang 외

The rocketing prosperity of large language models (LLMs) in recent years has boosted the prevalence of vision-language models (VLMs) in the medical sector. In our online medical consultation scenario, a doctor responds t…

Medical Visual Question AnsweringQuestion AnsweringVisual GroundingVisual Question Answering+1