Contrastive Cross-Modal Pre-Training: A General Strategy for Small Sample Medical Imaging
A key challenge in training neural networks for a given medical imaging task is often the difficulty of obtaining a sufficient number of manually labeled examples. In contrast, textual imaging reports, which are often readily available in medical records, contain rich but unstructured interpretations written by experts as part of standard clinical practice. We propose using these textual reports as a form of weak supervision to improve the image interpretation performance of a neural network without requiring additional manually labeled examples. We use an image-text matching task to train a feature extractor and then fine-tune it in a transfer learning setting for a supervised task using a small labeled dataset. The end result is a neural network that automatically interprets imagery without requiring textual reports during inference. This approach can be applied to any task for which text-image pairs are readily available. We evaluate our method on three classification tasks and find consistent performance improvements, reducing the need for labeled data by 67%-98%.
Code (0)
등록된 구현이 없습니다.
Tasks
Image ClassificationImage-text matchingText MatchingTransfer LearningSimilar Papers 제목 키워드 기반
Vision Language Pre-training by Contrastive Learning with Cross-Modal Similarity Regulation
Cross-modal contrastive learning in vision language pretraining (VLP) faces the challenge of (partial) false negatives. In this paper, we study this problem from the perspective of Mutual Information (MI) optimization. I…
Common Sense ReasoningContrastive LearningCross-Modal Contrastive Learning for Robust Reasoning in VQA
Multi-modal reasoning in visual question answering (VQA) has witnessed rapid progress recently. However, most reasoning models heavily rely on shortcuts learned from training data, which prevents their usage in challengi…
Contrastive LearningQuestion AnsweringTripletVisual Question Answering+1Exploiting Pseudo Image Captions for Multimodal Summarization
Cross-modal contrastive learning in vision language pretraining (VLP) faces the challenge of (partial) false negatives. In this paper, we study this problem from the perspective of Mutual Information (MI) optimization. I…
Common Sense ReasoningContrastive LearningImage CaptioningTurbo your multi-modal classification with contrastive learning
Contrastive learning has become one of the most impressive approaches for multi-modal representation learning. However, previous multi-modal works mainly focused on cross-modal understanding, ignoring in-modal contrastiv…
ClassificationContrastive LearningEmotion RecognitionMulti-modal Classification+4UMCL: Unimodal-generated Multimodal Contrastive Learning for Cross-compression-rate Deepfake Detection
In deepfake detection, the varying degrees of compression employed by social media platforms pose significant challenges for model generalization and reliability. Although existing methods have progressed from single-mod…
Contrastive LearningDeepFake Detection