paper-with-me

홈 › Papers

CLIP Adaptation by Intra-modal Overlap Reduction

2024-09-17 · Alexey Kravets, Vinay Namboodiri

Numerous methods have been proposed to adapt a pre-trained foundational CLIP model for few-shot classification. As CLIP is trained on a large corpus, it generalises well through adaptation to few-shot classification. In this work, we analyse the intra-modal overlap in image space in terms of embedding representation. Our analysis shows that, due to contrastive learning, embeddings from CLIP model exhibit high cosine similarity distribution overlap in the image space between paired and unpaired examples affecting the performance of few-shot training-free classification methods which rely on similarity in the image space for their predictions. To tackle intra-modal overlap we propose to train a lightweight adapter on a generic set of samples from the Google Open Images dataset demonstrating that this improves accuracy for few-shot training-free classification. We validate our contribution through extensive empirical analysis and demonstrate that reducing the intra-modal overlap leads to a) improved performance on a number of standard datasets, b) increased robustness to distribution shift and c) higher feature variance rendering the features more discriminative for downstream tasks.

📄 PDF Abstract BibTeX arXiv:2409.11338

Code (0)

등록된 구현이 없습니다.

Tasks

ClassificationContrastive Learning

Methods 이 논문이 사용한 방법론

SET Dynamic Sparse Training method where weight mask is updated randomly periodically
Adapter 설명 없음
CLIP Contrastive Language-Image Pre-training (CLIP), consisting of a simplified version of ConVIRT trained from scratch, is an efficient method of image representation learning…

Similar Papers 제목 키워드 기반

Connecting Multi-modal Contrastive Representations

2023-05-22 · NeurIPS 2023 11 · Zehan Wang, Yang Zhao, Xize Cheng, Haifeng Huang 외

Multi-modal Contrastive Representation learning aims to encode different modalities into a semantically aligned shared space. This paradigm shows remarkable generalization ability on numerous downstream tasks across vari…

3D Point Cloud ClassificationcounterfactualImage RetrievalPoint Cloud Classification+2

Domain Aligned CLIP for Few-shot Classification

2023-11-15 · Muhammad Waleed Gondal, Jochen Gast, Inigo Alonso Ruiz, Richard Droste 외

Large vision-language representation learning models like CLIP have demonstrated impressive performance for zero-shot transfer to downstream tasks while largely benefiting from inter-modal (image-text) alignment via cont…

BenchmarkingClassificationDomain Adaptationimage-classification+2

IsoCLIP: Decomposing CLIP Projectors for Efficient Intra-modal Alignment

2026-03-20 · Simone Magistri, Dipam Goswami, Marco Mistretta, Bartłomiej Twardowski 외 arxiv

Vision-Language Models like CLIP are extensively used for inter-modal tasks which involve both visual and text modalities. However, when the individual modality encoders are applied to inherently intra-modal tasks like i…

Image Retrieval

Mitigate the Gap: Investigating Approaches for Improving Cross-Modal Alignment in CLIP

2024-06-25 · Sedigheh Eslami, Gerard de Melo

Contrastive Language--Image Pre-training (CLIP) has manifested remarkable improvements in zero-shot classification and cross-modal vision-language tasks. Yet, from a geometrical point of view, the CLIP embedding space ha…

cross-modal alignmentImage Classificationtext similarityzero-shot-classification+2

Causal Disentanglement and Cross-Modal Alignment for Enhanced Few-Shot Learning

2025-08-05 · Tianjiao Jiang, Zhen Zhang, Yuhang Liu, Javen Qinfeng Shi arxiv

Few-shot learning (FSL) often requires effective adaptation of models using limited labeled data. However, most existing FSL methods rely on entangled representations, requiring the model to implicitly recover the unmixi…

Computational EfficiencyContrastive LearningFew-Shot Learning