An Empirical Study of Scaling Law for Scene Text Recognition
The laws of model size data volume computation and model performance have been extensively studied in the field of Natural Language Processing (NLP). However the scaling laws in Scene Text Recognition (STR) have not yet been investigated. To address this we conducted comprehensive studies that involved examining the correlations between performance and the scale of models data volume and computation in the field of text recognition. Conclusively the study demonstrates smooth power laws between performance and model size as well as training data volume when other influencing factors are held constant. Additionally we have constructed a large-scale dataset called REBU-Syn which comprises 6 million real samples and 18 million synthetic samples. Based on our scaling law and new dataset we have successfully trained a scene text recognition model achieving a new state-of-the-art on 6 common test benchmarks with a top-1 average accuracy of 97.42%. The models and dataset are publicly available at \href https://github.com/large-ocr-model/large-ocr-model.github.io large-ocr-model.github.io .
Code (1)
Tasks
Optical Character Recognition (OCR)Scene Text RecognitionSimilar Papers 제목 키워드 기반
An Empirical Study of Scaling Law for OCR
The laws of model size, data volume, computation and model performance have been extensively studied in the field of Natural Language Processing (NLP). However, the scaling laws in Optical Character Recognition (OCR) hav…
Optical Character RecognitionOptical Character Recognition (OCR)Scene Text RecognitionAccurate Scene Text Recognition with Efficient Model Scaling and Cloze Self-Distillation
Scaling architectures have been proven effective for improving Scene Text Recognition (STR), but the individual contribution of vision encoder and text decoder scaling remain under-explored. In this work, we present an i…
DecoderScene Text RecognitionAn Empirical Study of Remote Sensing Pretraining
Deep learning has largely reshaped remote sensing (RS) research for aerial image understanding and made a great success. Nevertheless, most of the existing deep models are initialized with the ImageNet pretrained weights…
Aerial Scene ClassificationBuilding change detection for remote sensing imagesChange DetectionChange detection for remote sensing images+4Learn to Augment: Joint Data Augmentation and Network Optimization for Text Recognition
Handwritten text and scene text suffer from various shapes and distorted patterns. Thus training a robust recognition model requires a large amount of data to cover diversity as much as possible. In contrast to data coll…
Data AugmentationDiversityImage AugmentationOn Model and Data Scaling for Skeleton-based Self-Supervised Gait Recognition
Gait recognition from video streams is a challenging problem in computer vision biometrics due to the subtle differences between gaits and numerous confounding factors. Recent advancements in self-supervised pretraining …
Gait Recognition