paper-with-me

홈 › Papers

Tip-Adapter: Training-free Adaption of CLIP for Few-shot Classification

2022-07-19 · Renrui Zhang, Zhang Wei, Rongyao Fang, Peng Gao, Kunchang Li, Jifeng Dai, Yu Qiao, Hongsheng Li

Contrastive Vision-Language Pre-training, known as CLIP, has provided a new paradigm for learning visual representations using large-scale image-text pairs. It shows impressive performance on downstream tasks by zero-shot knowledge transfer. To further enhance CLIP's adaption capability, existing methods proposed to fine-tune additional learnable modules, which significantly improves the few-shot performance but introduces extra training time and computational resources. In this paper, we propose a training-free adaption method for CLIP to conduct few-shot classification, termed as Tip-Adapter, which not only inherits the training-free advantage of zero-shot CLIP but also performs comparably to those training-required approaches. Tip-Adapter constructs the adapter via a key-value cache model from the few-shot training set, and updates the prior knowledge encoded in CLIP by feature retrieval. On top of that, the performance of Tip-Adapter can be further boosted to be state-of-the-art on ImageNet by fine-tuning the cache model for 10$\times$ fewer epochs than existing methods, which is both effective and efficient. We conduct extensive experiments of few-shot classification on 11 datasets to demonstrate the superiority of our proposed methods. Code is released at https://github.com/gaopengcuhk/Tip-Adapter.

📄 PDF Abstract BibTeX arXiv:2207.09519

Code (3)

gaopengcuhk/tip-adapter 공식 구현 pytorch
ArsenalCheng/Meta-Adapter pytorch
opengvlab/cafo pytorch

Tasks

RetrievalTransfer Learning

Methods 이 논문이 사용한 방법론

CLIP Contrastive Language-Image Pre-training (CLIP), consisting of a simplified version of ConVIRT trained from scratch, is an efficient method of image representation learning…
Adapter 설명 없음

Similar Papers 제목 키워드 기반

Hold-One-Shot-Out (HOSO) for Validation-Free Few-Shot CLIP Adapters

2026-03-04 · Chris Vorster, Mayug Maniparambil, Noel E. O'Connor, Noel Murphy 외 arxiv

In many CLIP adaptation methods, a blending ratio hyperparameter controls the trade-off between general pretrained CLIP knowledge and the limited, dataset-specific supervision from the few-shot cases. Most few-shot CLIP …

Tip-Adapter: Training-free CLIP-Adapter for Better Vision-Language Modeling

2021-11-06 · Renrui Zhang, Rongyao Fang, Wei zhang, Peng Gao 외

Contrastive Vision-Language Pre-training, known as CLIP, has provided a new paradigm for learning visual representations by using large-scale contrastive image-text pairs. It shows impressive performance on zero-shot kno…

Language ModelingLanguage ModellingTransfer Learning

Ta-Adapter: Enhancing few-shot CLIP with task-aware encoders

2024-04-29 · Pattern Recognition 153 (2024) 110559 2024 4 · Wenbo Zhang, Yifan Zhang, Yuyang Deng, Wenlong Zhang 외

Contrastive Language-Image Pre-training (CLIP) has shown impressive zero-shot transfer capabilities, but its potential for specific downstream tasks is not fully utilized. To further enhance CLIP’s few-shot capability fo…

Few-Shot Learningimage-classificationImage ClassificationPrompt Learning

Learning to Adapt Category Consistent Meta-Feature of CLIP for Few-Shot Classification

2024-07-08 · Jiaying Shi, Xuetong Xue, Shenghui Xu

The recent CLIP-based methods have shown promising zero-shot and few-shot performance on image classification tasks. Existing approaches such as CoOp and Tip-Adapter only focus on high-level visual features that are full…

Few-Shot Learningimage-classificationImage Classification

GP-Adapter: Gaussian Process CLIP-Adapter for Few-Shot Out-of-Distribution Detection

2026-06-05 · Taisei Saito, Koretaka Ogata, Takafumi Hiroi arxiv

We propose GP-Adapter, a training-free framework that augments CLIP (Contrastive Language-Image Pre-training) with Gaussian Process (GP) uncertainty modeling for few-shot classification and out-of-distribution (OOD) dete…

Out-of-Distribution Detection