Exploiting LMM-based knowledge for image classification tasks
In this paper we address image classification tasks leveraging knowledge encoded in Large Multimodal Models (LMMs). More specifically, we use the MiniGPT-4 model to extract semantic descriptions for the images, in a multimodal prompting fashion. In the current literature, vision language models such as CLIP, among other approaches, are utilized as feature extractors, using only the image encoder, for solving image classification tasks. In this paper, we propose to additionally use the text encoder to obtain the text embeddings corresponding to the MiniGPT-4-generated semantic descriptions. Thus, we use both the image and text embeddings for solving the image classification task. The experimental evaluation on three datasets validates the improved classification performance achieved by exploiting LMM-based knowledge.
Code (0)
등록된 구현이 없습니다.
Tasks
Classificationimage-classificationImage ClassificationMethods 이 논문이 사용한 방법론
Similar Papers 제목 키워드 기반
TAI++: Text as Image for Multi-Label Image Classification by Co-Learning Transferable Prompt
The recent introduction of prompt tuning based on pre-trained vision-language models has dramatically improved the performance of multi-label image classification. However, some existing strategies that have been explore…
Diversityimage-classificationImage ClassificationMulti-Label Image Classification+1Improving Scene Graph Classification by Exploiting Knowledge from Texts
Training scene graph classification models requires a large amount of annotated image data. Meanwhile, scene graphs represent relational knowledge that can be modeled with symbolic data from texts or knowledge graphs. Wh…
ClassificationGeneral ClassificationGraph ClassificationKnowledge Graphs+7Human Attention in Fine-grained Classification
The way humans attend to, process and classify a given image has the potential to vastly benefit the performance of deep learning models. Exploiting where humans are focusing can rectify models when they are deviating fr…
ClassificationDecision MakingFine-Grained Image ClassificationKnowledge-enhanced Visual-Language Pretraining for Computational Pathology
In this paper, we consider the problem of visual representation learning for computational pathology, by exploiting large-scale image-text pairs gathered from public resources, along with the domain-specific knowledge in…
Cross-Modal RetrievalLanguage ModelingLanguage ModellingRepresentation Learning+4Remote sensing image classification exploiting multiple kernel learning
We propose a strategy for land use classification which exploits Multiple Kernel Learning (MKL) to automatically determine a suitable combination of a set of features without requiring any heuristic knowledge about the c…
ClassificationGeneral Classificationimage-classificationImage Classification+1