Traditional Chinese Synthetic Datasets Verified with Labeled Data for Scene Text Recognition
Scene text recognition (STR) has been widely studied in academia and industry. Training a text recognition model often requires a large amount of labeled data, but data labeling can be difficult, expensive, or time-consuming, especially for Traditional Chinese text recognition. To the best of our knowledge, public datasets for Traditional Chinese text recognition are lacking. This paper presents a framework for a Traditional Chinese synthetic data engine which aims to improve text recognition model performance. We generated over 20 million synthetic data and collected over 7,000 manually labeled data TC-STR 7k-word as the benchmark. Experimental results show that a text recognition model can achieve much better accuracy either by training from scratch with our generated synthetic data or by further fine-tuning with TC-STR 7k-word.
Code (1)
Tasks
Scene Text RecognitionSimilar Papers 제목 키워드 기반
Data Augmentation for Low-resource Word Segmentation and POS Tagging of Ancient Chinese Texts
Automatic word segmentation and part-of-speech tagging of ancient books can help relevant researchers to study ancient texts. In recent years, pre-trained language models have achieved significant improvements on text pr…
Data AugmentationLanguage ModelingLanguage ModellingPart-Of-Speech Tagging+2PKUSEG: A Toolkit for Multi-Domain Chinese Word Segmentation
Chinese word segmentation (CWS) is a fundamental step of Chinese natural language processing. In this paper, we build a new toolkit, named PKUSEG, for multi-domain word segmentation. Unlike existing single-model toolkits…
Chinese Word SegmentationDomain AdaptationPOSPOS Tagging+1From Obstacles to Resources: Semi-supervised Learning Faces Synthetic Data Contamination
Semi-supervised learning (SSL) can improve model performance by leveraging unlabeled images, which can be collected from public image sources with low costs. In recent years, synthetic images have become increasingly com…
Training Data Synthesis with Difficulty Controlled Diffusion Model
Semi-supervised learning (SSL) can improve model performance by leveraging unlabeled images, which can be collected from public image sources with low costs. In recent years, synthetic images have become increasingly com…
modelTowards Automatic Generation of Entertaining Dialogues in Chinese Crosstalks
Crosstalk, also known by its Chinese name xiangsheng, is a traditional Chinese comedic performing art featuring jokes and funny dialogues, and one of China's most popular cultural elements. It is typically in the form of…
Dialogue GenerationTranslation