paper-with-me

홈 › Papers

Large-Scale Zero-Shot Image Classification from Rich and Diverse Textual Descriptions

2021-03-17 · EACL (LANTERN) 2021 4 · Sebastian Bujwid, Josephine Sullivan

We study the impact of using rich and diverse textual descriptions of classes for zero-shot learning (ZSL) on ImageNet. We create a new dataset ImageNet-Wiki that matches each ImageNet class to its corresponding Wikipedia article. We show that merely employing these Wikipedia articles as class descriptions yields much higher ZSL performance than prior works. Even a simple model using this type of auxiliary data outperforms state-of-the-art models that rely on standard features of word embedding encodings of class names. These results highlight the usefulness and importance of textual descriptions for ZSL, as well as the relative importance of auxiliary data type compared to algorithmic progress. Our experimental results also show that standard zero-shot learning approaches generalize poorly across categories of classes.

📄 PDF Abstract BibTeX arXiv:2103.09669

Code (0)

등록된 구현이 없습니다.

Tasks

ArticlesGeneral Classificationimage-classificationImage ClassificationZero-Shot Image ClassificationZero-Shot Learning

Similar Papers 제목 키워드 기반

I2MVFormer: Large Language Model Generated Multi-View Document Supervision for Zero-Shot Image Classification

2022-12-05 · CVPR 2023 1 · Muhammad Ferjad Naeem, Muhammad Gul Zain Ali Khan, Yongqin Xian, Muhammad Zeshan Afzal 외

Recent works have shown that unstructured text (documents) from online sources can serve as useful auxiliary information for zero-shot image classification. However, these methods require access to a high-quality source …

Classificationimage-classificationImage ClassificationLanguage Modeling+3

CARZero: Cross-Attention Alignment for Radiology Zero-Shot Classification

2024-02-27 · CVPR 2024 1 · Haoran Lai, Qingsong Yao, Zihang Jiang, Rongsheng Wang 외

The advancement of Zero-Shot Learning in the medical domain has been driven forward by using pre-trained models on large-scale image-text pairs, focusing on image-text alignment. However, existing methods primarily rely …

ClassificationDiagnosticLanguage ModellingLarge Language Model+3

Large-Scale Bidirectional Training for Zero-Shot Image Captioning

2022-11-13 · TaeHoon Kim, Mark Marsden, Pyunghwan Ahn, Sangyun Kim 외

When trained on large-scale datasets, image captioning models can understand the content of images from a general domain but often fail to generate accurate, detailed captions. To improve performance, pretraining-and-fin…

Image CaptioningKeyword Extraction

Improved Zero-Shot Classification by Adapting VLMs with Text Descriptions

2024-01-04 · CVPR 2024 1 · Oindrila Saha, Grant van Horn, Subhransu Maji

The zero-shot performance of existing vision-language models (VLMs) such as CLIP is limited by the availability of large-scale, aligned image and text datasets in specific domains. In this work, we leverage two complemen…

Fine-Grained Image Classificationimage-classificationImage Classificationzero-shot-classification+1

A ChatGPT Aided Explainable Framework for Zero-Shot Medical Image Diagnosis

2023-07-05 · Jiaxiang Liu, Tianxiang Hu, Yan Zhang, Xiaotang Gai 외

Zero-shot medical image classification is a critical process in real-world scenarios where we have limited access to all possible diseases or large-scale annotated data. It involves computing similarity scores between a …

Diagnosticimage-classificationImage ClassificationMedical Image Classification