paper-with-me

홈 › Papers

Towards Unbiased Multi-label Zero-Shot Learning with Pyramid and Semantic Attention

2022-03-07 · Ziming Liu, Song Guo, Jingcai Guo, Yuanyuan Xu, Fushuo Huo

Multi-label zero-shot learning extends conventional single-label zero-shot learning to a more realistic scenario that aims at recognizing multiple unseen labels of classes for each input sample. Existing works usually exploit attention mechanism to generate the correlation among different labels. However, most of them are usually biased on several major classes while neglect most of the minor classes with the same importance in input samples, and may thus result in overly diffused attention maps that cannot sufficiently cover minor classes. We argue that disregarding the connection between major and minor classes, i.e., correspond to the global and local information, respectively, is the cause of the problem. In this paper, we propose a novel framework of unbiased multi-label zero-shot learning, by considering various class-specific regions to calibrate the training process of the classifier. Specifically, Pyramid Feature Attention (PFA) is proposed to build the correlation between global and local information of samples to balance the presence of each class. Meanwhile, for the generated semantic representations of input samples, we propose Semantic Attention (SA) to strengthen the element-wise correlation among these vectors, which can encourage the coordinated representation of them. Extensive experiments on the large-scale multi-label zero-shot benchmarks NUS-WIDE and Open-Image demonstrate that the proposed method surpasses other representative methods by significant margins.

📄 PDF Abstract BibTeX arXiv:2203.03483

Code (0)

등록된 구현이 없습니다.

Tasks

Multi-label zero-shot learningZero-Shot Learning

Similar Papers 제목 키워드 기반

PRIMERA: Pyramid-based Masked Sentence Pre-training for Multi-document Summarization

2021-11-16 · ACL ARR November 2021 11 · Anonymous

We introduce PRIMERA, a pre-trained model for multi-document representation with a focus on summarization that reduces the need for dataset-specific architectures and large amounts of fine-tuning labeled data. PRIMERA us…

DecoderDocument SummarizationMulti-Document SummarizationSentence

PyraTok: Language-Aligned Pyramidal Tokenizer for Video Understanding and Generation

2026-01-22 · Onkar Susladkar, Tushar Prakash, Adheesh Juvekar, Kiet A. Nguyen 외 arxiv

Discrete video VAEs underpin modern text-to-video generation and video understanding systems, yet existing tokenizers typically learn visual codebooks at a single scale with limited vocabularies and shallow language supe…

Temporal Action LocalizationText-to-Video GenerationVideo ReconstructionVideo Segmentation

Learning unbiased zero-shot semantic segmentation networks via transductive transfer

2020-07-01 · Haiyang Liu, Yichen Wang, Jiayi Zhao, Guowu Yang 외

Semantic segmentation, which aims to acquire a detailed understanding of images, is an essential issue in computer vision. However, in practical scenarios, new categories that are different from the categories in trainin…

AttributePredictionSegmentationSemantic Segmentation+3

PRIMERA: Pyramid-based Masked Sentence Pre-training for Multi-document Summarization

2021-10-16 · ACL 2022 5 · Wen Xiao, Iz Beltagy, Giuseppe Carenini, Arman Cohan

We introduce PRIMERA, a pre-trained model for multi-document representation with a focus on summarization that reduces the need for dataset-specific architectures and large amounts of fine-tuning labeled data. PRIMERA us…

Abstractive Text SummarizationDecoderDocument SummarizationMulti-Document Summarization+2

Transductive Unbiased Embedding for Zero-Shot Learning

2018-03-30 · CVPR 2018 6 · Jie Song, Chengchao Shen, Yezhou Yang, Yang Liu 외

Most existing Zero-Shot Learning (ZSL) methods have the strong bias problem, in which instances of unseen (target) classes tend to be categorized as one of the seen (source) classes. So they yield poor performance after …

Transductive LearningZero-Shot Learning