paper-with-me

홈 › Papers

An Empirical Study of Scaling Law for Scene Text Recognition

2024-01-01 · CVPR 2024 1 · Miao Rang, Zhenni Bi, Chuanjian Liu, Yunhe Wang, Kai Han

The laws of model size data volume computation and model performance have been extensively studied in the field of Natural Language Processing (NLP). However the scaling laws in Scene Text Recognition (STR) have not yet been investigated. To address this we conducted comprehensive studies that involved examining the correlations between performance and the scale of models data volume and computation in the field of text recognition. Conclusively the study demonstrates smooth power laws between performance and model size as well as training data volume when other influencing factors are held constant. Additionally we have constructed a large-scale dataset called REBU-Syn which comprises 6 million real samples and 18 million synthetic samples. Based on our scaling law and new dataset we have successfully trained a scene text recognition model achieving a new state-of-the-art on 6 common test benchmarks with a top-1 average accuracy of 97.42%. The models and dataset are publicly available at \href https://github.com/large-ocr-model/large-ocr-model.github.io large-ocr-model.github.io .

📄 PDF Abstract BibTeX

Code (1)

large-ocr-model/large-ocr-model.github.io 공식 구현

Tasks

Optical Character Recognition (OCR)Scene Text Recognition

Similar Papers 제목 키워드 기반

An Empirical Study of Scaling Law for OCR

2023-12-29 · Miao Rang, Zhenni Bi, Chuanjian Liu, Yunhe Wang 외

The laws of model size, data volume, computation and model performance have been extensively studied in the field of Natural Language Processing (NLP). However, the scaling laws in Optical Character Recognition (OCR) hav…

Optical Character RecognitionOptical Character Recognition (OCR)Scene Text Recognition

Accurate Scene Text Recognition with Efficient Model Scaling and Cloze Self-Distillation

2025-03-20 · CVPR 2025 1 · Andrea Maracani, Savas Ozkan, Sijun Cho, Hyowon Kim 외

Scaling architectures have been proven effective for improving Scene Text Recognition (STR), but the individual contribution of vision encoder and text decoder scaling remain under-explored. In this work, we present an i…

DecoderScene Text Recognition

An Empirical Study of Remote Sensing Pretraining

2022-04-06 · Di Wang, Jing Zhang, Bo Du, Gui-Song Xia 외

Deep learning has largely reshaped remote sensing (RS) research for aerial image understanding and made a great success. Nevertheless, most of the existing deep models are initialized with the ImageNet pretrained weights…

Aerial Scene ClassificationBuilding change detection for remote sensing imagesChange DetectionChange detection for remote sensing images+4

Learn to Augment: Joint Data Augmentation and Network Optimization for Text Recognition

2020-03-14 · CVPR 2020 6 · Canjie Luo, Yuanzhi Zhu, Lianwen Jin, Yongpan Wang

Handwritten text and scene text suffer from various shapes and distorted patterns. Thus training a robust recognition model requires a large amount of data to cover diversity as much as possible. In contrast to data coll…

Data AugmentationDiversityImage Augmentation

On Model and Data Scaling for Skeleton-based Self-Supervised Gait Recognition

2025-04-10 · Adrian Cosma, Andy Cǎtrunǎ, Emilian Rǎdoi

Gait recognition from video streams is a challenging problem in computer vision biometrics due to the subtle differences between gaits and numerous confounding factors. Recent advancements in self-supervised pretraining …

Gait Recognition