paper-with-me

홈 › Papers

Transformer with Peak Suppression and Knowledge Guidance for Fine-grained Image Recognition

2021-07-14 · Xinda Liu, Lili Wang, Xiaoguang Han

Fine-grained image recognition is challenging because discriminative clues are usually fragmented, whether from a single image or multiple images. Despite their significant improvements, most existing methods still focus on the most discriminative parts from a single image, ignoring informative details in other regions and lacking consideration of clues from other associated images. In this paper, we analyze the difficulties of fine-grained image recognition from a new perspective and propose a transformer architecture with the peak suppression module and knowledge guidance module, which respects the diversification of discriminative features in a single image and the aggregation of discriminative clues among multiple images. Specifically, the peak suppression module first utilizes a linear projection to convert the input image into sequential tokens. It then blocks the token based on the attention response generated by the transformer encoder. This module penalizes the attention to the most discriminative parts in the feature learning process, therefore, enhancing the information exploitation of the neglected regions. The knowledge guidance module compares the image-based representation generated from the peak suppression module with the learnable knowledge embedding set to obtain the knowledge response coefficients. Afterwards, it formalizes the knowledge learning as a classification problem using response coefficients as the classification scores. Knowledge embeddings and image-based representations are updated during training so that the knowledge embedding includes discriminative clues for different images. Finally, we incorporate the acquired knowledge embeddings into the image-based representations as comprehensive representations, leading to significantly higher performance. Extensive evaluations on the six popular datasets demonstrate the advantage of the proposed method.

📄 PDF Abstract BibTeX arXiv:2107.06538

Code (0)

등록된 구현이 없습니다.

Tasks

Fine-Grained Image ClassificationFine-Grained Image Recognition

Similar Papers 제목 키워드 기반

Orthogonal Negative Guidance in Attention Feature Space for Text-to-Image Generation

2026-05-28 · Jungmin Ko, Jungwon Park, Jimyeong Kim, Changin Choi 외 arxiv

Text-to-image (T2I) models have become increasingly capable of generating high-quality images. Yet, enforcing the explicit absence of a specified object or attribute remains a fundamentally challenging problem. Existing …

Text-to-Image Generation

Artifact Correction for Echo-Planar Imaging at Low-Field and Ultra-Low-Field MRI

2026-05-25 · Sisi Qiao, Yilin Yu, Tiecheng Lin, Yuhao Liu 외 arxiv

Purpose: Echo-planar imaging (EPI) in low-field (LF) and ultra-low-field MRI (ULF) suffers from severe Nyquist ghost artifacts due to odd-even k-space misalignment. This study develops a reference-free artifact correctio…

A Hybrid Continuity Loss to Reduce Over-Suppression for Time-domain Target Speaker Extraction

2022-03-31 · Zexu Pan, Meng Ge, Haizhou Li

The speaker extraction algorithm extracts the target speech from a mixture speech containing interference speech and background noise. The extraction process sometimes over-suppresses the extracted target speech, which n…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)speech-recognitionSpeech Recognition+1

Impact of complex spatial population structure on early and long-term adaptation in rugged fitness landscapes

2024-09-25 · Richard Servajean, Arthur Alexandre, Anne-Florence Bitbol

We investigate the exploration of rugged fitness landscapes by spatially structured populations with demes on the nodes of a graph, connected by migrations. In the rare migration regime, we find that finite structures ca…

Personalized Speech Enhancement Without a Separate Speaker Embedding Model

2024-06-14 · Tanel Pärnamaa, Ando Saabas

Personalized speech enhancement (PSE) models can improve the audio quality of teleconferencing systems by adapting to the characteristics of a speaker's voice. However, most existing methods require a separate speaker em…

Speech Enhancement