Complementary Attributes: A New Clue to Zero-Shot Learning
Zero-shot learning (ZSL) aims to recognize unseen objects using disjoint seen objects via sharing attributes. The generalization performance of ZSL is governed by the attributes, which transfer semantic information from seen classes to unseen classes. To take full advantage of the knowledge transferred by attributes, in this paper, we introduce the notion of complementary attributes (CA), as a supplement to the original attributes, to enhance the semantic representation ability. Theoretical analyses demonstrate that complementary attributes can improve the PAC-style generalization bound of original ZSL model. Since the proposed CA focuses on enhancing the semantic representation, CA can be easily applied to any existing attribute-based ZSL methods, including the label-embedding strategy based ZSL (LEZSL) and the probability-prediction strategy based ZSL (PPZSL). In PPZSL, there is a strong assumption that all the attributes are independent of each other, which is arguably unrealistic in practice. To solve this problem, a novel rank aggregation framework is proposed to circumvent the assumption. Extensive experiments on five ZSL benchmark datasets and the large-scale ImageNet dataset demonstrate that the proposed complementary attributes and rank aggregation can significantly and robustly improve existing ZSL methods and achieve the state-of-the-art performance.
Code (0)
등록된 구현이 없습니다.
Tasks
AttributeStyle GeneralizationZero-Shot LearningSimilar Papers 제목 키워드 기반
Visual Clues: Bridging Vision and Language Foundations for Image Paragraph Captioning
People say, "A picture is worth a thousand words". Then how can we get the rich information out of the image? We argue that by using visual clues to bridge large pretrained vision foundation models and language models, w…
Image Paragraph CaptioningLanguage ModelingLanguage ModellingLarge Language ModelImproved Zero-Shot Classification by Adapting VLMs with Text Descriptions
The zero-shot performance of existing vision-language models (VLMs) such as CLIP is limited by the availability of large-scale, aligned image and text datasets in specific domains. In this work, we leverage two complemen…
Fine-Grained Image Classificationimage-classificationImage Classificationzero-shot-classification+1Italian Crossword Generator: Enhancing Education through Interactive Word Puzzles
Educational crosswords offer numerous benefits for students, including increased engagement, improved understanding, critical thinking, and memory retention. Creating high-quality educational crosswords can be challengin…
Few-Shot LearningZero-Shot LearningMulti-Modal Prototypes for Open-World Semantic Segmentation
In semantic segmentation, generalizing a visual system to both seen categories and novel categories at inference time has always been practically valuable yet challenging. To enable such functionality, existing methods m…
SegmentationSemantic SegmentationFrom Zero-shot Learning to Conventional Supervised Classification: Unseen Visual Data Synthesis
Robust object recognition systems usually rely on powerful feature extraction mechanisms from a large number of real images. However, in many realistic applications, collecting sufficient images for ever-growing new clas…
General ClassificationObject RecognitionZero-Shot Learning