paper-with-me

홈 › Papers

What If We Only Use Real Datasets for Scene Text Recognition? Toward Scene Text Recognition With Fewer Labels

2021-03-07 · CVPR 2021 1 · Jeonghun Baek, Yusuke Matsui, Kiyoharu Aizawa

Scene text recognition (STR) task has a common practice: All state-of-the-art STR models are trained on large synthetic data. In contrast to this practice, training STR models only on fewer real labels (STR with fewer labels) is important when we have to train STR models without synthetic data: for handwritten or artistic texts that are difficult to generate synthetically and for languages other than English for which we do not always have synthetic data. However, there has been implicit common knowledge that training STR models on real data is nearly impossible because real data is insufficient. We consider that this common knowledge has obstructed the study of STR with fewer labels. In this work, we would like to reactivate STR with fewer labels by disproving the common knowledge. We consolidate recently accumulated public real data and show that we can train STR models satisfactorily only with real labeled data. Subsequently, we find simple data augmentation to fully exploit real data. Furthermore, we improve the models by collecting unlabeled data and introducing semi- and self-supervised methods. As a result, we obtain a competitive model to state-of-the-art methods. To the best of our knowledge, this is the first study that 1) shows sufficient performance by only using real labels and 2) introduces semi- and self-supervised methods into STR with fewer labels. Our code and data are available: https://github.com/ku21fan/STR-Fewer-Labels

📄 PDF Abstract BibTeX arXiv:2103.04400

Code (1)

ku21fan/STR-Fewer-Labels 공식 구현 pytorch

Tasks

Data AugmentationScene Text Recognition

Similar Papers 제목 키워드 기반

MTRNet: A Generic Scene Text Eraser

2019-03-11 · Osman Tursun, Rui Zeng, Simon Denman, Sabesan Sivapalan 외

Text removal algorithms have been proposed for uni-lingual scripts with regular shapes and layouts. However, to the best of our knowledge, a generic text removal method which is able to remove all or user-specified text …

Curved Text DetectionText Detection

What Images Cannot Say: Language-Guided Olfactory Representation Learning

2026-07-07 · Eleftherios Tsonis, Xi Wang, Vicky Kalogeiton arxiv

Images tell us what a scene looks like, but rarely what it would feel like to be there. While recent datasets pair visual scenes with electronic-nose measurements, aligning smell signals with images remains challenging b…

Representation LearningText Retrieval

What Is Wrong with Synthetic Data for Scene Text Recognition? A Strong Synthetic Engine with Diverse Simulations and Self-Evolution

2026-02-06 · Xingsong Ye, Yongkun Du, JiaXin Zhang, Chen Li 외 arxiv

Large-scale and categorical-balanced text data is essential for training effective Scene Text Recognition (STR) models, which is hard to achieve when collecting real data. Synthetic data offers a cost-effective and perfe…

Scene Text Recognition

Choose What You Need: Disentangled Representation Learning for Scene Text Recognition Removal and Editing

2024-01-01 · CVPR 2024 1 · Boqiang Zhang, Hongtao Xie, Zuan Gao, Yuxin Wang

Scene text images contain not only style information (font background) but also content information (character texture). Different scene text tasks need different information but previous representation learning meth…

DecoderRepresentation LearningScene Text Recognition

Choose What You Need: Disentangled Representation Learning for Scene Text Recognition, Removal and Editing

2024-05-07 · Boqiang Zhang, Hongtao Xie, Zuan Gao, Yuxin Wang

Scene text images contain not only style information (font, background) but also content information (character, texture). Different scene text tasks need different information, but previous representation learning metho…

DecoderRepresentation LearningScene Text Recognition