Region-based Discriminative Feature Pooling for Scene Text Recognition
We present a new feature representation method for scene text recognition problem, particularly focusing on improving scene character recognition. Many existing methods rely on Histogram of Oriented Gradient (HOG) or part-based models, which do not span the feature space well for characters in natural scene images, especially given large variation in fonts with cluttered backgrounds. In this work, we propose a discriminative feature pooling method that automatically learns the most informative sub-regions of each scene character within a multi-class classification framework, whereas each sub-region seamlessly integrates a set of low-level image features through integral images. The proposed feature representation is compact, computationally efficient, and able to effectively model distinctive spatial structures of each individual character class. Extensive experiments conducted on challenging datasets (Chars74K, ICDAR'03, ICDAR'11, SVT) show that our method significantly outperforms existing methods on scene character classification and scene text recognition tasks.
Code (0)
등록된 구현이 없습니다.
Tasks
General ClassificationMulti-class ClassificationScene Text RecognitionSimilar Papers 제목 키워드 기반
Learning Important Spatial Pooling Regions for Scene Classification
We address the false response influence problem when learning and applying discriminative parts to construct the mid-level representation in scene classification. It is often caused by the complexity of latent image stru…
ClassificationGeneral ClassificationScene ClassificationHarvesting Discriminative Meta Objects with Deep CNN Features for Scene Classification
Recent work on scene classification still makes use of generic CNN features in a rudimentary manner. In this ICCV 2015 paper, we present a novel pipeline built upon deep CNN features to harvest discriminative visual obje…
ClusteringGeneral ClassificationRegion ProposalScene Classification+1A Novel Dual-pooling Attention Module for UAV Vehicle Re-identification
Vehicle re-identification (Re-ID) involves identifying the same vehicle captured by other cameras, given a vehicle image. It plays a crucial role in the development of safe cities and smart cities. With the rapid growth …
Single Particle AnalysisTripletVehicle Re-IdentificationContext-aware Attentional Pooling (CAP) for Fine-grained Visual Classification
Deep convolutional neural networks (CNNs) have shown a strong ability in mining discriminative object pose and parts information for image recognition. For fine-grained recognition, context-aware rich feature representat…
Fine-Grained Image ClassificationGeneral ClassificationInformativenessObjectLearning to Recognize Actions on Objects in Egocentric Video with Attention Dictionaries
We present EgoACO, a deep neural architecture for video action recognition that learns to pool action-context-object descriptors from frame level features by leveraging the verb-noun structure of action labels in egocent…
Action RecognitionObjectTemporal Action Localization