paper-with-me

홈 › Papers

Learning to Discover Multi-Class Attentional Regions for Multi-Label Image Recognition

2020-07-03 · Bin-Bin Gao, Hong-Yu Zhou

Multi-label image recognition is a practical and challenging task compared to single-label image classification. However, previous works may be suboptimal because of a great number of object proposals or complex attentional region generation modules. In this paper, we propose a simple but efficient two-stream framework to recognize multi-category objects from global image to local regions, similar to how human beings perceive objects. To bridge the gap between global and local streams, we propose a multi-class attentional region module which aims to make the number of attentional regions as small as possible and keep the diversity of these regions as high as possible. Our method can efficiently and effectively recognize multi-class objects with an affordable computation cost and a parameter-free region localization module. Over three benchmarks on multi-label image classification, we create new state-of-the-art results with a single model only using image semantics without label dependency. In addition, the effectiveness of the proposed method is extensively demonstrated under different factors such as global pooling strategy, input size and network architecture. Code has been made available at~\url{https://github.com/gaobb/MCAR}.

📄 PDF Abstract BibTeX arXiv:2007.01755

Code (1)

gaobb/MCAR 공식 구현 pytorch

Tasks

Diversityimage-classificationMulti-Label ClassificationMulti-Label Image ClassificationMulti-Label Image Recognition

Similar Papers 제목 키워드 기반

Multi-label Image Recognition by Recurrently Discovering Attentional Regions

2017-11-08 · ICCV 2017 10 · Zhouxia Wang, Tianshui Chen, Guanbin Li, Ruijia Xu 외

This paper proposes a novel deep architecture to address multi-label image recognition, a fundamental and practical task towards general visual understanding. Current solutions for this task usually rely on an extra step…

General Classificationimage-classificationImage ClassificationMulti-Label Image Classification+2

Recurrent Attentional Reinforcement Learning for Multi-label Image Recognition

2017-12-20 · Tianshui Chen, Zhouxia Wang, Guanbin Li, Liang Lin

Recognizing multiple labels of images is a fundamental but challenging task in computer vision, and remarkable progress has been attained by localizing semantic-aware image regions and predicting their labels with deep c…

Multi-Label Image Recognitionreinforcement-learningReinforcement LearningReinforcement Learning (RL)

AttnGAN: Fine-Grained Text to Image Generation with Attentional Generative Adversarial Networks

2017-11-28 · CVPR 2018 6 · Tao Xu, Pengchuan Zhang, Qiuyuan Huang, Han Zhang 외

In this paper, we propose an Attentional Generative Adversarial Network (AttnGAN) that allows attention-driven, multi-stage refinement for fine-grained text-to-image generation. With a novel attentional generative networ…

Generative Adversarial NetworkImage GenerationImage-text matchingText Matching+2

Visual Explanations from Hadamard Product in Multimodal Deep Networks

2017-12-18 · Jin-Hwa Kim, Byoung-Tak Zhang

The visual explanation of learned representation of models helps to understand the fundamentals of learning. The attentional models of previous works used to visualize the attended regions over an image or text using the…

Question AnsweringVisual Question AnsweringVisual Question Answering (VQA)

Marginalized Average Attentional Network for Weakly-Supervised Learning

2019-05-21 · ICLR 2019 5 · Yuan Yuan, Yueming Lyu, Xi Shen, Ivor W. Tsang 외

In weakly-supervised temporal action localization, previous works have failed to locate dense and integral regions for each entire action due to the overestimation of the most salient regions. To alleviate this issue, we…

Action LocalizationTemporal Action LocalizationWeakly Supervised Action LocalizationWeakly-supervised Learning+1