paper-with-me

홈 › Papers

Dual Cross-Attention Learning for Fine-Grained Visual Categorization and Object Re-Identification

2022-05-04 · CVPR 2022 1 · Haowei Zhu, Wenjing Ke, Dong Li, Ji Liu, Lu Tian, Yi Shan

Recently, self-attention mechanisms have shown impressive performance in various NLP and CV tasks, which can help capture sequential characteristics and derive global information. In this work, we explore how to extend self-attention modules to better learn subtle feature embeddings for recognizing fine-grained objects, e.g., different bird species or person identities. To this end, we propose a dual cross-attention learning (DCAL) algorithm to coordinate with self-attention learning. First, we propose global-local cross-attention (GLCA) to enhance the interactions between global images and local high-response regions, which can help reinforce the spatial-wise discriminative clues for recognition. Second, we propose pair-wise cross-attention (PWCA) to establish the interactions between image pairs. PWCA can regularize the attention learning of an image by treating another image as distractor and will be removed during inference. We observe that DCAL can reduce misleading attentions and diffuse the attention response to discover more complementary parts for recognition. We conduct extensive evaluations on fine-grained visual categorization and object re-identification. Experiments demonstrate that DCAL performs on par with state-of-the-art methods and consistently improves multiple self-attention baselines, e.g., surpassing DeiT-Tiny and ViT-Base by 2.8% and 2.4% mAP on MSMT17, respectively.

📄 PDF Abstract BibTeX arXiv:2205.02151

Code (0)

등록된 구현이 없습니다.

Tasks

Fine-Grained Image ClassificationFine-Grained Visual Categorization

Similar Papers 제목 키워드 기반

Dual Capsule Attention Mask Network with Mutual Learning for Visual Question Answering

2022-10-01 · COLING 2022 10 · Weidong Tian, Haodong Li, Zhong-Qiu Zhao

A Visual Question Answering (VQA) model processes images and questions simultaneously with rich semantic information. The attention mechanism can highlight fine-grained features with critical information, thus ensuring t…

Question AnsweringVisual Question AnsweringVisual Question Answering (VQA)

Fine-Grained Attention for Weakly Supervised Object Localization

2021-04-11 · Junghyo Sohn, Eunjin Jeon, Wonsik Jung, Eunsong Kang 외

Although recent advances in deep learning accelerated an improvement in a weakly supervised object localization (WSOL) task, there are still challenges to identify the entire body of an object, rather than only discrimin…

ObjectObject LocalizationWeakly-Supervised Object Localization

ConCA: Concentration-Aware Channel Attention for Fine-Grained Visual Recognition

2026-08-31 · Yu-Sheng Liu, Yu-Chen Tung arxiv

Lightweight channel attention mechanisms are widely used in image classification, yet their effectiveness in fine-grained visual recognition (FGVR) remains limited. Most modules summarize each channel by global average p…

Fine-Grained Visual RecognitionImage Classification

Hierarchical Vision-Language Interaction for Facial Action Unit Detection

2026-02-16 · Yong Li, Yi Ren, Yizhe Zhang, Wenhua Zhang 외 arxiv

Facial Action Unit (AU) detection seeks to recognize subtle facial muscle activations as defined by the Facial Action Coding System (FACS). A primary challenge w.r.t AU detection is the effective learning of discriminati…

Facial Action Unit DetectionRepresentation Learning

Cross-layer Attention Network for Fine-grained Visual Categorization

2022-10-17 · Ranran Huang, Yu Wang, Huazhong Yang

Learning discriminative representations for subtle localized details plays a significant role in Fine-grained Visual Categorization (FGVC). Compared to previous attention-based works, our work does not explicitly define …

Fine-Grained Visual Categorization