Generalised Zero-Shot Learning with Domain Classification in a Joint Semantic and Visual Space
Generalised zero-shot learning (GZSL) is a classification problem where the learning stage relies on a set of seen visual classes and the inference stage aims to identify both the seen visual classes and a new set of unseen visual classes. Critically, both the learning and inference stages can leverage a semantic representation that is available for the seen and unseen classes. Most state-of-the-art GZSL approaches rely on a mapping between latent visual and semantic spaces without considering if a particular sample belongs to the set of seen or unseen classes. In this paper, we propose a novel GZSL method that learns a joint latent representation that combines both visual and semantic information. This mitigates the need for learning a mapping between the two spaces. Our method also introduces a domain classification that estimates whether a sample belongs to a seen or an unseen class. Our classifier then combines a class discriminator with this domain classifier with the goal of reducing the natural bias that GZSL approaches have toward the seen classes. Experiments show that our method achieves state-of-the-art results in terms of harmonic mean, the area under the seen and unseen curve and unseen classification accuracy on public GZSL benchmark data sets. Our code will be available upon acceptance of this paper.
Code (0)
등록된 구현이 없습니다.
Tasks
domain classificationGeneral ClassificationZero-Shot LearningSimilar Papers 제목 키워드 기반
Generalised Zero-Shot Learning with a Classifier Ensemble over Multi-Modal Embedding Spaces
Generalised zero-shot learning (GZSL) methods aim to classify previously seen and unseen visual classes by leveraging the semantic information of those classes. In the context of GZSL, semantic information is non-visual …
General ClassificationModel SelectionZero-Shot LearningTemporal and cross-modal attention for audio-visual zero-shot learning
Audio-visual generalised zero-shot learning for video classification requires understanding the relations between the audio and visual information in order to be able to recognise samples from novel, previously unseen cl…
GZSL Video ClassificationVideo ClassificationZero-Shot LearningAudio-visual Generalised Zero-shot Learning with Cross-modal Attention and Language
Learning to classify video data from classes not included in the training data, i.e. video-based zero-shot learning, is challenging. We conjecture that the natural alignment between the audio and visual modalities in vid…
GZSL Video ClassificationZero-Shot LearningZSL Video ClassificationMetaAudio: A Few-Shot Audio Classification Benchmark
Currently available benchmarks for few-shot learning (machine learning with few training examples) are limited in the domains they cover, primarily focusing on image classification. This work aims to alleviate this relia…
Audio ClassificationClassificationFew-Shot Audio ClassificationFew-Shot Learning+3Image-free Classifier Injection for Zero-Shot Classification
Zero-shot learning models achieve remarkable results on image classification for samples from classes that were not seen during training. However, such models must be trained from scratch with specialised methods: theref…
ClassificationDecoderimage-classificationImage Classification+2