paper-with-me

홈 › Papers

CLIP-IT: CLIP-based Pairing for Histology Images Classification

2025-04-22 · Banafsheh Karimian, Giulia Avanzato, Soufian Belharbi, Luke McCaffrey, Mohammadhadi Shateri, Eric Granger

Multimodal learning has shown significant promise for improving medical image analysis by integrating information from complementary data sources. This is widely employed for training vision-language models (VLMs) for cancer detection based on histology images and text reports. However, one of the main limitations in training these VLMs is the requirement for large paired datasets, raising concerns over privacy, and data collection, annotation, and maintenance costs. To address this challenge, we introduce CLIP-IT method to train a vision backbone model to classify histology images by pairing them with privileged textual information from an external source. At first, the modality pairing step relies on a CLIP-based model to match histology images with semantically relevant textual report data from external sources, creating an augmented multimodal dataset without the need for manually paired samples. Then, we propose a multimodal training procedure that distills the knowledge from the paired text modality to the unimodal image classifier for enhanced performance without the need for the textual data during inference. A parameter-efficient fine-tuning method is used to efficiently address the misalignment between the main (image) and paired (text) modalities. During inference, the improved unimodal histology classifier is used, with only minimal additional computational complexity. Our experiments on challenging PCAM, CRC, and BACH histology image datasets show that CLIP-IT can provide a cost-effective approach to leverage privileged textual information and outperform unimodal classifiers for histology.

📄 PDF Abstract BibTeX arXiv:2504.16181

Code (1)

BanafshehKarimian/ModalityPairing 공식 구현 pytorch

Tasks

Medical Image Analysisparameter-efficient fine-tuning

Similar Papers 제목 키워드 기반

CLIP is Strong Enough to Fight Back: Test-time Counterattacks towards Zero-shot Adversarial Robustness of CLIP

2025-03-05 · CVPR 2025 1 · Songlong Xing, Zhengyu Zhao, Nicu Sebe

Despite its prevalent use in image-text matching tasks in a zero-shot manner, CLIP has been shown to be highly vulnerable to adversarial perturbations added onto images. Recent studies propose to finetune the vision enco…

Adversarial RobustnessImage-text matchingText Matching

Image declipping with deep networks

2018-11-15 · Shachar Honig, Michael Werman

We present a deep network to recover pixel values lost to clipping. The clipped area of the image is typically a uniform area of minimum or maximum brightness, losing image detail and color fidelity. The degree to which …

Image Declipping

MoDE: CLIP Data Experts via Clustering

2024-04-24 · CVPR 2024 1 · Jiawei Ma, Po-Yao Huang, Saining Xie, Shang-Wen Li 외

The success of contrastive language-image pretraining (CLIP) relies on the supervision from the pairing between images and captions, which tends to be noisy in web-crawled data. We present Mixture of Data Experts (MoDE) …

Clusteringimage-classificationImage ClassificationZero-Shot Image Classification

Learn "No" to Say "Yes" Better: Improving Vision-Language Models via Negations

2024-03-29 · Jaisidh Singh, Ishaan Shrivastava, Mayank Vatsa, Richa Singh 외

Existing vision-language models (VLMs) treat text descriptions as a unit, confusing individual concepts in a prompt and impairing visual semantic matching and reasoning. An important aspect of reasoning in logic and lang…

image-classificationImage ClassificationZero-Shot Image Classification

SpectralCLIP: Preventing Artifacts in Text-Guided Style Transfer from a Spectral Perspective

2023-03-16 · Zipeng Xu, Songlong Xing, Enver Sangineto, Nicu Sebe

Owing to the power of vision-language foundation models, e.g., CLIP, the area of image synthesis has seen recent important advances. Particularly, for style transfer, CLIP enables transferring more general and abstract s…

Image GenerationStyle Transfer