MT: Multi-Perspective Feature Learning Network for Scene Text Detection
Text detection, the key technology for understanding scene text, has become an attractive research topic. For detecting various scene texts, researchers propose plenty of detectors with different advantages: detection-based models enjoy fast detection speed, and segmentation-based algorithms are not limited by text shapes. However, for most intelligent systems, the detector needs to detect arbitrary-shaped texts with high speed and accuracy simultaneously. Thus, in this study, we design an efficient pipeline named as MT, which can detect adhesive arbitrary-shaped texts with only a single binary mask in the inference stage. This paper presents the contributions on three aspects: (1) a light-weight detection framework is designed to speed up the inference process while keeping high detection accuracy; (2) a multi-perspective feature module is proposed to learn more discriminative representations to segment the mask accurately; (3) a multi-factor constraints IoU minimization loss is introduced for training the proposed model. The effectiveness of MT is evaluated on four real-world scene text datasets, and it surpasses all the state-of-the-art competitors to a large extent.
Code (0)
등록된 구현이 없습니다.
Tasks
Scene Text DetectionText DetectionSimilar Papers 제목 키워드 기반
Fused Text Segmentation Networks for Multi-oriented Scene Text Detection
In this paper, we introduce a novel end-end framework for multi-oriented scene text detection from an instance-aware semantic segmentation perspective. We present Fused Text Segmentation Networks, which combine multi-lev…
Multi-Oriented Scene Text Detectionobject-detectionObject DetectionRegion Proposal+5Perspective-aware Convolution for Monocular 3D Object Detection
Monocular 3D object detection is a crucial and challenging task for autonomous driving vehicle, while it uses only a single camera image to infer 3D objects in the scene. To address the difficulty of predicting depth usi…
3D Object DetectionAutonomous DrivingMonocular 3D Object DetectionObject+2Enhancing Multi-View Pedestrian Detection Through Generalized 3D Feature Pulling
The main challenge in multi-view pedestrian detection is integrating view-specific features into a unified space for comprehensive end-to-end perception. Prior multi-view detection methods have focused on projecting pers…
multi-view detectionMultiview DetectionPedestrian DetectionvalidEnhanced Multi-View Pedestrian Detection Using Probabilistic Occupancy Volume
Occlusion poses a significant challenge in pedestrian detection from a single view. To address this, multi-view detection systems have been utilized to aggregate information from multiple perspectives. Recent advances in…
3D Reconstructionmulti-view detectionPedestrian DetectionIncepText: A New Inception-Text Module with Deformable PSROI Pooling for Multi-Oriented Scene Text Detection
Incidental scene text detection, especially for multi-oriented text regions, is one of the most challenging tasks in many computer vision applications. Different from the common object detection task, scene text often su…
Multi-Oriented Scene Text Detectionobject-detectionObject DetectionOptical Character Recognition (OCR)+2