paper-with-me

Papers

Fine-grained Image Classification and Retrieval by Combining Visual and Locally Pooled Textual Features

2020-01-14 · Andres Mafla, Sounak Dey, Ali Furkan Biten, Lluis Gomez, Dimosthenis Karatzas

Text contained in an image carries high-level semantics that can be exploited to achieve richer image understanding. In particular, the mere presence of text provides strong guiding content that should be employed to tackle a diversity of computer vision tasks such as image retrieval, fine-grained classification, and visual question answering. In this paper, we address the problem of fine-grained classification and image retrieval by leveraging textual information along with visual cues to comprehend the existing intrinsic relation between the two modalities. The novelty of the proposed model consists of the usage of a PHOC descriptor to construct a bag of textual words along with a Fisher Vector Encoding that captures the morphology of text. This approach provides a stronger multimodal representation for this task and as our experiments demonstrate, it achieves state-of-the-art results on two different tasks, fine-grained classification and image retrieval.

📄 PDF Abstract BibTeX arXiv:2001.04732

Code (2)

DreadPiratePsyopus/Fine_Grained_Clf 공식 구현 pytorch
AndresPMD/Fine_Grained_Clf pytorch

Tasks

ClassificationDiversityFine-Grained Image ClassificationGeneral Classificationimage-classificationImage ClassificationImage RetrievalQuestion AnsweringRetrievalVisual Question AnsweringVisual Question Answering (VQA)

Similar Papers 제목 키워드 기반

You'll Never Walk Alone: A Sketch and Text Duet for Fine-Grained Image Retrieval

2024-03-12 · CVPR 2024 1 · Subhadeep Koley, Ayan Kumar Bhunia, Aneeshan Sain, Pinaki Nath Chowdhury 외

Two primary input modalities prevail in image retrieval: sketch and text. While text is widely used for inter-category retrieval tasks, sketches have been established as the sole preferred modality for fine-grained image…

AttributeImage RetrievalRetrieval

Integrating Scene Text and Visual Appearance for Fine-Grained Image Classification

2017-04-15 · Xiang Bai, Mingkun Yang, Pengyuan Lyu, Yongchao Xu 외

Text in natural images contains rich semantics that are often highly relevant to objects or scene. In this paper, we focus on the problem of fully exploiting scene text for visual understanding. The main idea is combinin…

ClassificationFine-Grained Image ClassificationGeneral Classificationimage-classification+2

Selective Convolutional Descriptor Aggregation for Fine-Grained Image Retrieval

2016-04-18 · Xiu-Shen Wei, Jian-Hao Luo, Jianxin Wu, Zhi-Hua Zhou

Deep convolutional neural network models pre-trained for the ImageNet classification task have been successfully adopted to tasks in other domains, such as texture description and object proposal generation, but these ta…

Image RetrievalObject Proposal GenerationRetrieval

Fine-graind Image Classification via Combining Vision and Language

2017-04-10 · Xiangteng He, Yuxin Peng

Fine-grained image classification is a challenging task due to the large intra-class variance and small inter-class variance, aiming at recognizing hundreds of sub-categories belonging to the same basic-level category. M…

AttributeClassificationFine-Grained Image ClassificationGeneral Classification+2

Multi-Modal Reasoning Graph for Scene-Text Based Fine-Grained Image Classification and Retrieval

2020-09-21 · Andres Mafla, Sounak Dey, Ali Furkan Biten, Lluis Gomez 외

Scene text instances found in natural images carry explicit semantic information that can provide important cues to solve a wide array of computer vision problems. In this paper, we focus on leveraging multi-modal conten…

Fine-Grained Image ClassificationGeneral Classificationimage-classificationImage Classification+2