An Erudite Fine-Grained Visual Classification Model
Current fine-grained visual classification (FGVC) models are isolated. In practice, we first need to identify the coarse-grained label of an object, then select the corresponding FGVC model for recognition. This hinders the application of the FGVC algorithm in real-life scenarios. In this paper, we propose an erudite FGVC model jointly trained by several different datasets, which can efficiently and accurately predict an object's fine-grained label across the combined label space. We found through a pilot study that positive and negative transfers co-occur when different datasets are mixed for training, i.e., the knowledge from other datasets is not always useful. Therefore, we first propose a feature disentanglement module and a feature re-fusion module to reduce negative transfer and boost positive transfer between different datasets. In detail, we reduce negative transfer by decoupling the deep features through many dataset-specific feature extractors. Subsequently, these are channel-wise re-fused to facilitate positive transfer. Finally, we propose a meta-learning based dataset-agnostic spatial attention layer to take full advantage of the multi-dataset training data, given that localisation is dataset-agnostic between different datasets. Experimental results across 11 different mixed-datasets built on four different FGVC datasets demonstrate the effectiveness of the proposed method. Furthermore, the proposed method can be easily combined with existing FGVC methods to obtain state-of-the-art results.
Code (0)
등록된 구현이 없습니다.
Tasks
ClassificationDisentanglementFine-Grained Image ClassificationMeta-LearningmodelSimilar Papers 제목 키워드 기반
ERUDITE: Human-in-the-Loop IoT for an Adaptive Personalized Learning System
Thanks to the rapid growth in wearable technologies and recent advancement in machine learning and signal processing, monitoring complex human contexts becomes feasible, paving the way to develop human-in-the-loop IoT sy…
Learning TheoryA Novel Plug-in Module for Fine-Grained Visual Classification
Visual classification can be divided into coarse-grained and fine-grained classification. Coarse-grained classification represents categories with a large degree of dissimilarity, such as the classification of cats and d…
ClassificationFine-Grained Image ClassificationFine-Grained Image RecognitionUnderstanding the Fine-Grained Knowledge Capabilities of Vision-Language Models
Vision-language models (VLMs) have made substantial progress across a wide range of visual question answering benchmarks, spanning visual reasoning, document understanding, and multimodal dialogue. These improvements are…
Visual Question AnsweringImage ClassificationVisual ReasoningGeneralizable Whole Slide Image Classification with Fine-Grained Visual-Semantic Interaction
Whole Slide Image (WSI) classification is often formulated as a Multiple Instance Learning (MIL) problem. Recently, Vision-Language Models (VLMs) have demonstrated remarkable performance in WSI classification. However, e…
image-classificationImage ClassificationLanguage ModellingLarge Language Model+2Coarse2Fine: A Two-stage Training Method for Fine-grained Visual Classification
Small inter-class and large intra-class variations are the main challenges in fine-grained visual classification. Objects from different classes share visually similar structures and objects in the same class can have di…
Fine-Grained Image ClassificationGeneral ClassificationVocal Bursts Valence Prediction