paper-with-me

홈 › Papers

Chinese Text Recognition with A Pre-Trained CLIP-Like Model Through Image-IDS Aligning

2023-09-03 · ICCV 2023 1 · Haiyang Yu, Xiaocong Wang, Bin Li, xiangyang xue

Scene text recognition has been studied for decades due to its broad applications. However, despite Chinese characters possessing different characteristics from Latin characters, such as complex inner structures and large categories, few methods have been proposed for Chinese Text Recognition (CTR). Particularly, the characteristic of large categories poses challenges in dealing with zero-shot and few-shot Chinese characters. In this paper, inspired by the way humans recognize Chinese texts, we propose a two-stage framework for CTR. Firstly, we pre-train a CLIP-like model through aligning printed character images and Ideographic Description Sequences (IDS). This pre-training stage simulates humans recognizing Chinese characters and obtains the canonical representation of each character. Subsequently, the learned representations are employed to supervise the CTR model, such that traditional single-character recognition can be improved to text-line recognition through image-IDS matching. To evaluate the effectiveness of the proposed method, we conduct extensive experiments on both Chinese character recognition (CCR) and CTR. The experimental results demonstrate that the proposed method performs best in CCR and outperforms previous methods in most scenarios of the CTR benchmark. It is worth noting that the proposed method can recognize zero-shot Chinese characters in text images without fine-tuning, whereas previous methods require fine-tuning when new classes appear. The code is available at https://github.com/FudanVI/FudanOCR/tree/main/image-ids-CTR.

📄 PDF Abstract BibTeX arXiv:2309.01083

Code (1)

fudanvi/fudanocr 공식 구현 pytorch

Tasks

Scene Text Recognition

Similar Papers 제목 키워드 기반

Chinese CLIP: Contrastive Vision-Language Pretraining in Chinese

2022-11-02 · An Yang, Junshu Pan, Junyang Lin, Rui Men 외

The tremendous success of CLIP (Radford et al., 2021) has promoted the research and application of contrastive learning for vision-language pretraining. In this work, we construct a large-scale dataset of image-text pair…

Contrastive Learningimage-classificationImage ClassificationImage Retrieval+7

FG-CLIP 2: A Bilingual Fine-grained Vision-Language Alignment Model

2025-10-13 · Chunyu Xie, Bin Wang, Fanjing Kong, Jincheng Li 외 arxiv

Fine-grained vision-language understanding requires precise alignment between visual content and linguistic descriptions, a capability that remains limited in current models, particularly in non-English settings. While m…

A Progressive Framework of Vision-language Knowledge Distillation and Alignment for Multilingual Scene

2024-04-17 · Wenbo Zhang, Yifan Zhang, Jianfeng Lin, Binqiang Huang 외

Pre-trained vision-language (V-L) models such as CLIP have shown excellent performance in many downstream cross-modal tasks. However, most of them are only applicable to the English context. Subsequent research has focus…

image-classificationImage ClassificationKnowledge DistillationLanguage Modelling+1

A CNN Based Scene Chinese Text Recognition Algorithm With Synthetic Data Engine

2016-04-07 · Xiaohang Ren, Kai Chen, Jun Sun

Scene text recognition plays an important role in many computer vision applications. The small size of available public available scene text datasets is the main challenge when training a text recognition CNN model. In t…

Scene Text Recognition

Beyond Cropped Regions: New Benchmark and Corresponding Baseline for Chinese Scene Text Retrieval in Diverse Layouts

2025-06-05 · Gengluo Li, Huawen Shen, Yu Zhou

Chinese scene text retrieval is a practical task that aims to search for images containing visual instances of a Chinese query text. This task is extremely challenging because Chinese text often features complex and dive…

RetrievalText Retrieval