Large-Scale Zero-Shot Image Classification from Rich and Diverse Textual Descriptions
We study the impact of using rich and diverse textual descriptions of classes for zero-shot learning (ZSL) on ImageNet. We create a new dataset ImageNet-Wiki that matches each ImageNet class to its corresponding Wikipedia article. We show that merely employing these Wikipedia articles as class descriptions yields much higher ZSL performance than prior works. Even a simple model using this type of auxiliary data outperforms state-of-the-art models that rely on standard features of word embedding encodings of class names. These results highlight the usefulness and importance of textual descriptions for ZSL, as well as the relative importance of auxiliary data type compared to algorithmic progress. Our experimental results also show that standard zero-shot learning approaches generalize poorly across categories of classes.
Code (0)
등록된 구현이 없습니다.
Tasks
ArticlesGeneral Classificationimage-classificationImage ClassificationZero-Shot Image ClassificationZero-Shot LearningSimilar Papers 제목 키워드 기반
I2MVFormer: Large Language Model Generated Multi-View Document Supervision for Zero-Shot Image Classification
Recent works have shown that unstructured text (documents) from online sources can serve as useful auxiliary information for zero-shot image classification. However, these methods require access to a high-quality source …
Classificationimage-classificationImage ClassificationLanguage Modeling+3CARZero: Cross-Attention Alignment for Radiology Zero-Shot Classification
The advancement of Zero-Shot Learning in the medical domain has been driven forward by using pre-trained models on large-scale image-text pairs, focusing on image-text alignment. However, existing methods primarily rely …
ClassificationDiagnosticLanguage ModellingLarge Language Model+3Large-Scale Bidirectional Training for Zero-Shot Image Captioning
When trained on large-scale datasets, image captioning models can understand the content of images from a general domain but often fail to generate accurate, detailed captions. To improve performance, pretraining-and-fin…
Image CaptioningKeyword ExtractionImproved Zero-Shot Classification by Adapting VLMs with Text Descriptions
The zero-shot performance of existing vision-language models (VLMs) such as CLIP is limited by the availability of large-scale, aligned image and text datasets in specific domains. In this work, we leverage two complemen…
Fine-Grained Image Classificationimage-classificationImage Classificationzero-shot-classification+1A ChatGPT Aided Explainable Framework for Zero-Shot Medical Image Diagnosis
Zero-shot medical image classification is a critical process in real-world scenarios where we have limited access to all possible diseases or large-scale annotated data. It involves computing similarity scores between a …
Diagnosticimage-classificationImage ClassificationMedical Image Classification