paper-with-me

Papers

Integrating Scene Text and Visual Appearance for Fine-Grained Image Classification

2017-04-15 · Xiang Bai, Mingkun Yang, Pengyuan Lyu, Yongchao Xu, Jiebo Luo

Text in natural images contains rich semantics that are often highly relevant to objects or scene. In this paper, we focus on the problem of fully exploiting scene text for visual understanding. The main idea is combining word representations and deep visual features into a globally trainable deep convolutional neural network. First, the recognized words are obtained by a scene text reading system. Then, we combine the word embedding of the recognized words and the deep visual features into a single representation, which is optimized by a convolutional neural network for fine-grained image classification. In our framework, the attention mechanism is adopted to reveal the relevance between each recognized word and the given image, which further enhances the recognition performance. We have performed experiments on two datasets: Con-Text dataset and Drink Bottle dataset, that are proposed for fine-grained classification of business places and drink bottles, respectively. The experimental results consistently demonstrate that the proposed method combining textual and visual cues significantly outperforms classification with only visual representations. Moreover, we have shown that the learned representation improves the retrieval performance on the drink bottle images by a large margin, making it potentially useful in product search.

📄 PDF Abstract BibTeX arXiv:1704.04613

Code (0)

등록된 구현이 없습니다.

Tasks

ClassificationFine-Grained Image ClassificationGeneral Classificationimage-classificationImage ClassificationRetrieval

Similar Papers 제목 키워드 기반

Integrating Language-Derived Appearance Elements with Visual Cues in Pedestrian Detection

2023-11-02 · Sungjune Park, Hyunjun Kim, Yong Man Ro

Large language models (LLMs) have shown their capabilities in understanding contextual and semantic information regarding knowledge of instance appearances. In this paper, we introduce a novel approach to utilize the str…

Pedestrian Detection

IllumiCraft: Unified Geometry and Illumination Diffusion for Controllable Video Generation

2025-06-03 · Yuanze Lin, Yi-Wen Chen, Yi-Hsuan Tsai, Ronald Clark 외

Although diffusion-based models can generate high-quality and high-resolution video sequences from textual or image inputs, they lack explicit integration of geometric cues when controlling scene lighting and visual appe…

3D geometryVideo Generation

State and Scene Enhanced Prototypes for Weakly Supervised Open-Vocabulary Object Detection

2025-11-22 · Jiaying Zhou, Qingchao Chen arxiv

Open-Vocabulary Object Detection (OVOD) aims to generalize object recognition to novel categories, while Weakly Supervised OVOD (WS-OVOD) extends this by combining box-level annotations with image-level labels. Despite r…

Object RecognitionObject Detection

3D-Aware Scene Manipulation via Inverse Graphics

2018-08-28 · NeurIPS 2018 12 · Shunyu Yao, Tzu Ming Harry Hsu, Jun-Yan Zhu, Jiajun Wu 외

We aim to obtain an interpretable, expressive, and disentangled scene representation that contains comprehensive structural and textural information for each object. Previous scene representations learned by neural netwo…

DecoderDisentanglementObject

Linguistics-aware Masked Image Modeling for Self-supervised Scene Text Recognition

2025-03-24 · CVPR 2025 1 · Yifei Zhang, Chang Liu, Jin Wei, Xiaomeng Yang 외

Text images are unique in their dual nature, encompassing both visual and linguistic information. The visual component encompasses structural and appearance-based features, while the linguistic dimension incorporates con…

Contrastive LearningScene Text RecognitionSelf-Supervised Learningself-supervised scene text recognition