Fine-grained Image Classification and Retrieval by Combining Visual and Locally Pooled Textual Features
Text contained in an image carries high-level semantics that can be exploited to achieve richer image understanding. In particular, the mere presence of text provides strong guiding content that should be employed to tackle a diversity of computer vision tasks such as image retrieval, fine-grained classification, and visual question answering. In this paper, we address the problem of fine-grained classification and image retrieval by leveraging textual information along with visual cues to comprehend the existing intrinsic relation between the two modalities. The novelty of the proposed model consists of the usage of a PHOC descriptor to construct a bag of textual words along with a Fisher Vector Encoding that captures the morphology of text. This approach provides a stronger multimodal representation for this task and as our experiments demonstrate, it achieves state-of-the-art results on two different tasks, fine-grained classification and image retrieval.
Code (2)
Tasks
ClassificationDiversityFine-Grained Image ClassificationGeneral Classificationimage-classificationImage ClassificationImage RetrievalQuestion AnsweringRetrievalVisual Question AnsweringVisual Question Answering (VQA)Similar Papers 제목 키워드 기반
You'll Never Walk Alone: A Sketch and Text Duet for Fine-Grained Image Retrieval
Two primary input modalities prevail in image retrieval: sketch and text. While text is widely used for inter-category retrieval tasks, sketches have been established as the sole preferred modality for fine-grained image…
AttributeImage RetrievalRetrievalIntegrating Scene Text and Visual Appearance for Fine-Grained Image Classification
Text in natural images contains rich semantics that are often highly relevant to objects or scene. In this paper, we focus on the problem of fully exploiting scene text for visual understanding. The main idea is combinin…
ClassificationFine-Grained Image ClassificationGeneral Classificationimage-classification+2Selective Convolutional Descriptor Aggregation for Fine-Grained Image Retrieval
Deep convolutional neural network models pre-trained for the ImageNet classification task have been successfully adopted to tasks in other domains, such as texture description and object proposal generation, but these ta…
Image RetrievalObject Proposal GenerationRetrievalFine-graind Image Classification via Combining Vision and Language
Fine-grained image classification is a challenging task due to the large intra-class variance and small inter-class variance, aiming at recognizing hundreds of sub-categories belonging to the same basic-level category. M…
AttributeClassificationFine-Grained Image ClassificationGeneral Classification+2Multi-Modal Reasoning Graph for Scene-Text Based Fine-Grained Image Classification and Retrieval
Scene text instances found in natural images carry explicit semantic information that can provide important cues to solve a wide array of computer vision problems. In this paper, we focus on leveraging multi-modal conten…
Fine-Grained Image ClassificationGeneral Classificationimage-classificationImage Classification+2