Strokelets: A Learned Multi-Scale Representation for Scene Text Recognition
Driven by the wide range of applications, scene text detection and recognition have become active research topics in computer vision. Though extensively studied, localizing and reading text in uncontrolled environments remain extremely challenging, due to various interference factors. In this paper, we propose a novel multi-scale representation for scene text recognition. This representation consists of a set of detectable primitives, termed as strokelets, which capture the essential substructures of characters at different granularities. Strokelets possess four distinctive advantages: (1) Usability: automatically learned from bounding box labels; (2) Robustness: insensitive to interference factors; (3) Generality: applicable to variant languages; and (4) Expressivity: effective at describing characters. Extensive experiments on standard benchmarks verify the advantages of strokelets and demonstrate the effectiveness of the proposed algorithm for text recognition.
Code (0)
등록된 구현이 없습니다.
Tasks
Scene Text DetectionScene Text RecognitionText DetectionSimilar Papers 제목 키워드 기반
Visuomotor Understanding for Representation Learning of Driving Scenes
Dashboard cameras capture a tremendous amount of driving scene video each day. These videos are purposefully coupled with vehicle sensing data, such as from the speedometer and inertial sensors, providing an additional s…
Optical Flow EstimationRepresentation LearningScene UnderstandingSemantic SegmentationShape2Scene: 3D Scene Representation Learning Through Pre-training on Shape Data
Current 3D self-supervised learning methods of 3D scenes face a data desert issue, resulting from the time-consuming and expensive collecting process of 3D scene data. Conversely, 3D shape datasets are easier to collect.…
3D Object Detection3D Semantic Segmentationobject-detectionObject Detection+4Block-NeRF: Scalable Large Scene Neural View Synthesis
We present Block-NeRF, a variant of Neural Radiance Fields that can represent large-scale environments. Specifically, we demonstrate that when scaling NeRF to render city-scale scenes spanning multiple blocks, it is vita…
NeRFSceneDreamer: Unbounded 3D Scene Generation from 2D Image Collections
In this work, we present SceneDreamer, an unconditional generative model for unbounded 3D scenes, which synthesizes large-scale 3D landscapes from random noise. Our framework is learned from in-the-wild 2D image collecti…
Scene GenerationMultimodal Scale Consistency and Awareness for Monocular Self-Supervised Depth Estimation
Dense depth estimation is essential to scene-understanding for autonomous driving. However, recent self-supervised approaches on monocular videos suffer from scale-inconsistency across long sequences. Utilizing data from…
Autonomous DrivingDepth EstimationMonocular Depth EstimationScene Understanding