Deep Reinforced Attention Learning for Quality-Aware Visual Recognition
In this paper, we build upon the weakly-supervised generation mechanism of intermediate attention maps in any convolutional neural networks and disclose the effectiveness of attention modules more straightforwardly to fully exploit their potential. Given an existing neural network equipped with arbitrary attention modules, we introduce a meta critic network to evaluate the quality of attention maps in the main network. Due to the discreteness of our designed reward, the proposed learning method is arranged in a reinforcement learning setting, where the attention actors and recurrent critics are alternately optimized to provide instant critique and revision for the temporary attention representation, hence coined as Deep REinforced Attention Learning (DREAL). It could be applied universally to network architectures with different types of attention modules and promotes their expressive ability by maximizing the relative gain of the final recognition performance arising from each individual attention module, as demonstrated by extensive experiments on both category and instance recognition benchmarks.
Code (0)
등록된 구현이 없습니다.
Similar Papers 제목 키워드 기반
Contrastively-reinforced Attention Convolutional Neural Network for Fine-grained Image Recognition
Fine-grained visual classification is inherently challenging because of its inter-class similarity and intra-class variance. However, by contrasting the images with same/different labels, a human can instinctively notice…
ClassificationFine-Grained Image ClassificationFine-Grained Image RecognitionSemantic Reinforced Attention Learning for Visual Place Recognition
Large-scale visual place recognition (VPR) is inherently challenging because not all visual cues in the image are beneficial to the task. In order to highlight the task-relevant visual cues in the feature embedding, the …
Visual Place RecognitionProgressive Modality Reinforcement for Human Multimodal Emotion Recognition From Unaligned Multimodal Sequences
Human multimodal emotion recognition involves time-series data of different modalities, such as natural language, visual motions, and acoustic behaviors. Due to the variable sampling rates for sequences from differen…
Emotion RecognitionMultimodal Emotion RecognitionTime SeriesTime Series AnalysisAttentional Pyramid Pooling of Salient Visual Residuals for Place Recognition
The core of visual place recognition (VPR) lies in how to identify task-relevant visual cues and embed them into discriminative representations. Focusing on these two points, we propose a novel encoding strategy name…
Visual Place RecognitionEVLP:Learning Unified Embodied Vision-Language Planner with Reinforced Supervised Fine-Tuning
In complex embodied long-horizon manipulation tasks, effective task decomposition and execution require synergistic integration of textual logical reasoning and visual-spatial imagination to ensure efficient and accurate…
multimodal generationLogical Reasoning