Geometry-Aware Scene Text Detection With Instance Transformation Network
Localizing text in the wild is challenging in the situations of complicated geometric layout of the targets like random orientation and large aspect ratio. In this paper, we propose a geometry-aware modeling approach tailored for scene text representation with an end-to-end learning scheme. In our approach, a novel Instance Transformation Network (ITN) is presented to learn the geometry-aware representation encoding the unique geometric configurations of scene text instances with in-network transformation embedding, resulting in a robust and elegant framework to detect words or text lines at one pass. An end-to-end multi-task learning strategy with transformation regression, text/non-text classification and coordinate regression is adopted in the ITN. Experiments on the benchmark datasets demonstrate the effectiveness of the proposed approach in detecting scene text in various geometric configurations.
Code (0)
등록된 구현이 없습니다.
Tasks
General ClassificationMulti-Task LearningregressionScene Text Detectiontext-classificationText ClassificationText DetectionSimilar Papers 제목 키워드 기반
Geometry Normalization Networks for Accurate Scene Text Detection
Large geometry (e.g., orientation) variances are the key challenges in the scene text detection. In this work, we first conduct experiments to investigate the capacity of networks for learning geometry variances on detec…
Scene Text DetectionText DetectionBoosting Instance Awareness via Cross-View Correlation with 4D Radar and Camera for 3D Object Detection
4D millimeter-wave radar has emerged as a promising sensing modality for autonomous driving due to its robustness and affordability. However, its sparse and weak geometric cues make reliable instance activation difficult…
Scene Understanding3D Object DetectionAutonomous DrivingDisARM: Displacement Aware Relation Module for 3D Detection
We introduce Displacement Aware Relation Module (DisARM), a novel neural network module for enhancing the performance of 3D object detection in point cloud scenes. The core idea of our method is that contextual informati…
3D Object Detectionobject-detectionObject DetectionRelationGA-DAN: Geometry-Aware Domain Adaptation Network for Scene Text Detection and Recognition
Recent adversarial learning research has achieved very impressive progress for modelling cross-domain data shifts in appearance space but its counterpart in modelling cross-domain shifts in geometry space lags far behind…
Domain AdaptationScene Text DetectionText DetectionContext-Nav: Context-Driven Exploration and Viewpoint-Aware 3D Spatial Reasoning for Instance Navigation
Text-goal instance navigation (TGIN) asks an agent to resolve a single, free-form description into actions that reach the correct object instance among same-category distractors. We present \textit{Context-Nav}, which el…
Spatial Reasoning