paper-with-me

Papers

Traditional Chinese Synthetic Datasets Verified with Labeled Data for Scene Text Recognition

2021-11-26 · Yi-Chang Chen, Yu-Chuan Chang, Yen-Cheng Chang, Yi-Ren Yeh

Scene text recognition (STR) has been widely studied in academia and industry. Training a text recognition model often requires a large amount of labeled data, but data labeling can be difficult, expensive, or time-consuming, especially for Traditional Chinese text recognition. To the best of our knowledge, public datasets for Traditional Chinese text recognition are lacking. This paper presents a framework for a Traditional Chinese synthetic data engine which aims to improve text recognition model performance. We generated over 20 million synthetic data and collected over 7,000 manually labeled data TC-STR 7k-word as the benchmark. Experimental results show that a text recognition model can achieve much better accuracy either by training from scratch with our generated synthetic data or by further fine-tuning with TC-STR 7k-word.

📄 PDF Abstract BibTeX arXiv:2111.13327

Code (1)

gitycc/traditional-chinese-text-recogn-dataset 공식 구현

Tasks

Scene Text Recognition

Similar Papers 제목 키워드 기반

Data Augmentation for Low-resource Word Segmentation and POS Tagging of Ancient Chinese Texts

2022-06-01 · LT4HALA (LREC) 2022 6 · Yutong Shen, Jiahuan Li, ShuJian Huang, Yi Zhou 외

Automatic word segmentation and part-of-speech tagging of ancient books can help relevant researchers to study ancient texts. In recent years, pre-trained language models have achieved significant improvements on text pr…

Data AugmentationLanguage ModelingLanguage ModellingPart-Of-Speech Tagging+2

PKUSEG: A Toolkit for Multi-Domain Chinese Word Segmentation

2019-06-27 · Ruixuan Luo, Jingjing Xu, Yi Zhang, Zhiyuan Zhang 외

Chinese word segmentation (CWS) is a fundamental step of Chinese natural language processing. In this paper, we build a new toolkit, named PKUSEG, for multi-domain word segmentation. Unlike existing single-model toolkits…

Chinese Word SegmentationDomain AdaptationPOSPOS Tagging+1

From Obstacles to Resources: Semi-supervised Learning Faces Synthetic Data Contamination

2024-05-27 · Zerun Wang, Jiafeng Mao, Liuyu Xiang, Toshihiko Yamasaki

Semi-supervised learning (SSL) can improve model performance by leveraging unlabeled images, which can be collected from public image sources with low costs. In recent years, synthetic images have become increasingly com…

Training Data Synthesis with Difficulty Controlled Diffusion Model

2024-11-27 · Zerun Wang, Jiafeng Mao, Xueting Wang, Toshihiko Yamasaki

Semi-supervised learning (SSL) can improve model performance by leveraging unlabeled images, which can be collected from public image sources with low costs. In recent years, synthetic images have become increasingly com…

model

Towards Automatic Generation of Entertaining Dialogues in Chinese Crosstalks

2017-11-01 · Shikang Du, Xiaojun Wan, Yajie Ye

Crosstalk, also known by its Chinese name xiangsheng, is a traditional Chinese comedic performing art featuring jokes and funny dialogues, and one of China's most popular cultural elements. It is typically in the form of…

Dialogue GenerationTranslation