ELoPE: Fine-Grained Visual Classification with Efficient Localization, Pooling and Embedding
The task of fine-grained visual classification (FGVC) deals with classification problems that display a small inter-class variance such as distinguishing between different bird species or car models. State-of-the-art approaches typically tackle this problem by integrating an elaborate attention mechanism or (part-) localization method into a standard convolutional neural network (CNN). Also in this work the aim is to enhance the performance of a backbone CNN such as ResNet by including three efficient and lightweight components specifically designed for FGVC. This is achieved by using global k-max pooling, a discriminative embedding layer trained by optimizing class means and an efficient bounding box estimator that only needs class labels for training. The resulting model achieves new best state-of-the-art recognition accuracies on the Stanford cars and FGVC-Aircraft datasets.
Code (1)
Tasks
Fine-Grained Image ClassificationGeneral ClassificationMethods 이 논문이 사용한 방법론
Similar Papers 제목 키워드 기반
Fine-Grained Visual Classification with Efficient End-to-end Localization
The term fine-grained visual classification (FGVC) refers to classification tasks where the classes are very similar and the classification model needs to be able to find subtle differences to make the correct prediction…
ClassificationFine-Grained Image ClassificationGeneral ClassificationPindrop it! Audio and Visual Deepfake Countermeasures for Robust Detection and Fine Grained-Localization
The field of visual and audio generation is burgeoning with new state-of-the-art methods. This rapid proliferation of new techniques underscores the need for robust solutions for detecting synthetic content in videos. In…
Video ClassificationAudio GenerationWeakly-supervised Object Localization for Few-shot Learning and Fine-grained Few-shot Learning
Few-shot learning (FSL) aims to learn novel visual categories from very few samples, which is a challenging problem in real-world applications. Many methods of few-shot classification work well on general images to learn…
ClassificationFew-Shot LearningGeneral ClassificationObject Localization+1Language-guided Hierarchical Fine-grained Image Forgery Detection and Localization
Differences in forgery attributes of images generated in CNN-synthesized and image-editing domains are large, and such differences make a unified image forgery detection and localization (IFDL) challenging. To this end, …
AttributeImage Forgery DetectionRepresentation LearningGRE Suite: Geo-localization Inference via Fine-Tuned Vision-Language Models and Enhanced Reasoning Chains
Recent advances in Visual Language Models (VLMs) have demonstrated exceptional performance in visual reasoning tasks. However, geo-localization presents unique challenges, requiring the extraction of multigranular visual…
geo-localizationVisual ReasoningWorld Knowledge