paper-with-me

홈 › Papers

SuS-X: Training-Free Name-Only Transfer of Vision-Language Models

2022-11-28 · ICCV 2023 1 · Vishaal Udandarao, Ankush Gupta, Samuel Albanie

Contrastive Language-Image Pre-training (CLIP) has emerged as a simple yet effective way to train large-scale vision-language models. CLIP demonstrates impressive zero-shot classification and retrieval on diverse downstream tasks. However, to leverage its full potential, fine-tuning still appears to be necessary. Fine-tuning the entire CLIP model can be resource-intensive and unstable. Moreover, recent methods that aim to circumvent this need for fine-tuning still require access to images from the target distribution. In this paper, we pursue a different approach and explore the regime of training-free "name-only transfer" in which the only knowledge we possess about the downstream task comprises the names of downstream target categories. We propose a novel method, SuS-X, consisting of two key building blocks -- SuS and TIP-X, that requires neither intensive fine-tuning nor costly labelled data. SuS-X achieves state-of-the-art zero-shot classification results on 19 benchmark datasets. We further show the utility of TIP-X in the training-free few-shot setting, where we again achieve state-of-the-art results over strong training-free baselines. Code is available at https://github.com/vishaal27/SuS-X.

📄 PDF Abstract BibTeX arXiv:2211.16198

Code (2)

vishaal27/sus-x 공식 구현 pytorch
bethgelab/frequency_determines_performance pytorch

Tasks

Retrievalzero-shot-classificationZero-Shot Learning

Methods 이 논문이 사용한 방법론

CLIP Contrastive Language-Image Pre-training (CLIP), consisting of a simplified version of ConVIRT trained from scratch, is an efficient method of image representation learning…

Similar Papers 제목 키워드 기반

When Does Synthetic CT Transfer? A Label-Free Donor/Host Diagnostic for Medical Vision-Language Model Routing on Real Lung CT

2026-06-28 · Fakrul Islam Tushar arxiv

A synthetic measurement of model competence is useful only if it survives the move to real data, yet the real labels that would verify it are exactly what medical imaging lacks. We ask whether transfer can be predicted i…

Generalized Radiograph Representation Learning via Cross-supervision between Images and Free-text Radiology Reports

2021-11-04 · Hong-Yu Zhou, Xiaoyu Chen, Yinghao Zhang, Ruibang Luo 외

Pre-training lays the foundation for recent successes in radiograph analysis supported by deep learning. It learns transferable image representations by conducting large-scale fully-supervised or self-supervised learning…

Representation LearningSelf-Supervised LearningTransfer Learning

Polygon-free: Unconstrained Scene Text Detection with Box Annotations

2020-11-26 · Weijia Wu, Enze Xie, Ruimao Zhang, Wenhai Wang 외

Although a polygon is a more accurate representation than an upright bounding box for text detection, the annotations of polygons are extremely expensive and challenging. Unlike existing works that employ fully-supervise…

Scene Text DetectionText Detection

Cap2Det: Learning to Amplify Weak Caption Supervision for Object Detection

2019-07-23 · ICCV 2019 10 · Keren Ye, Mingda Zhang, Adriana Kovashka, Wei Li 외

Learning to localize and name object instances is a fundamental problem in vision, but state-of-the-art approaches rely on expensive bounding box supervision. While weakly supervised detection (WSOD) methods relax the ne…

Objectobject-detectionObject Detection

RadJEPA: Radiology Encoder for Chest X-Rays via Joint Embedding Predictive Architecture

2026-01-22 · Anas Anwarul Haq Khan, Mariam Husain, Pratik Jalan, Kshitij Jadhav arxiv

Vision-language pretraining has driven progress in medical image representation learning, but it depends on paired image-text data and can inherit reporting bias from clinical narratives. We study whether language-free p…

Representation LearningSemantic Segmentation