Attend and Rectify: a Gated Attention Mechanism for Fine-Grained Recovery
We propose a novel attention mechanism to enhance Convolutional Neural Networks for fine-grained recognition. It learns to attend to lower-level feature activations without requiring part annotations and uses these activations to update and rectify the output likelihood distribution. In contrast to other approaches, the proposed mechanism is modular, architecture-independent and efficient both in terms of parameters and computation required. Experiments show that networks augmented with our approach systematically improve their classification accuracy and become more robust to clutter. As a result, Wide Residual Networks augmented with our proposal surpasses the state of the art classification accuracies in CIFAR-10, the Adience gender recognition task, Stanford dogs, and UEC Food-100.
Code (1)
Tasks
ClassificationGeneral ClassificationImage ClassificationSimilar Papers 제목 키워드 기반
Not All Attention Is Needed: Gated Attention Network for Sequence Data
Although deep neural networks generally have fixed network structures, the concept of dynamic mechanism has drawn more and more attention in recent years. Attention mechanisms compute input-dependent dynamic attention we…
AllSentencetext-classificationText ClassificationCo-attending Regions and Detections with Multi-modal Multiplicative Embedding for VQA
Recently, the Visual Question Answering (VQA) task has gained increasing attention in artificial intelligence. Existing VQA methods mainly adopt the visual attention mechanism to associate the input question with corresp…
FormQuestion AnsweringVisual Question AnsweringVisual Question Answering (VQA)Co-attending Free-form Regions and Detections with Multi-modal Multiplicative Feature Embedding for Visual Question Answering
Recently, the Visual Question Answering (VQA) task has gained increasing attention in artificial intelligence. Existing VQA methods mainly adopt the visual attention mechanism to associate the input question with corresp…
FormVisual Question AnsweringVisual Question Answering (VQA)Gated Sparse Attention: Combining Computational Efficiency with Training Stability for Long-Context Language Models
The computational burden of attention in long-context language models has motivated two largely independent lines of work: sparse attention mechanisms that reduce complexity by attending to selected tokens, and gated att…
Computational EfficiencySystem 2 Attention (is something you might need too)
Soft attention in Transformer-based Large Language Models (LLMs) is susceptible to incorporating irrelevant information from the context into its latent representations, which adversely affects next token generations. To…
Math