paper-with-me

홈 › Papers

I2C2W: Image-to-Character-to-Word Transformers for Accurate Scene Text Recognition

2021-05-18 · Chuhui Xue, Jiaxing Huang, Wenqing Zhang, Shijian Lu, Changhu Wang, Song Bai

Leveraging the advances of natural language processing, most recent scene text recognizers adopt an encoder-decoder architecture where text images are first converted to representative features and then a sequence of characters via `sequential decoding'. However, scene text images suffer from rich noises of different sources such as complex background and geometric distortions which often confuse the decoder and lead to incorrect alignment of visual features at noisy decoding time steps. This paper presents I2C2W, a novel scene text recognition technique that is tolerant to geometric and photometric degradation by decomposing scene text recognition into two inter-connected tasks. The first task focuses on image-to-character (I2C) mapping which detects a set of character candidates from images based on different alignments of visual features in an non-sequential way. The second task tackles character-to-word (C2W) mapping which recognizes scene text by decoding words from the detected character candidates. The direct learning from character semantics (instead of noisy image features) corrects falsely detected character candidates effectively which improves the final text recognition accuracy greatly. Extensive experiments over nine public datasets show that the proposed I2C2W outperforms the state-of-the-art by large margins for challenging scene text datasets with various curvature and perspective distortions. It also achieves very competitive recognition performance over multiple normal scene text datasets.

📄 PDF Abstract BibTeX arXiv:2105.08383

Code (0)

등록된 구현이 없습니다.

Tasks

DecoderScene Text Recognition

Similar Papers 제목 키워드 기반

Confidence-aware Non-repetitive Multimodal Transformers for TextCaps

2020-12-07 · Zhaokai Wang, Renda Bao, Qi Wu, Si Liu

When describing an image, reading text in the visual scene is crucial to understand the key information. Recent work explores the TextCaps task, i.e. image captioning with reading Optical Character Recognition (OCR) toke…

Image CaptioningOptical Character RecognitionOptical Character Recognition (OCR)

Soft-PHOC Descriptor for End-to-End Word Spotting in Egocentric Scene Images

2018-09-04 · Dena Bazazian, Dimosthenis Karatzas, Andrew D. Bagdanov

Word spotting in natural scene images has many applications in scene understanding and visual assistance. In this paper we propose a technique to create and exploit an intermediate representation of images based on text …

AttributeDynamic Time WarpingScene Understanding

Layout Agnostic Scene Text Image Synthesis with Diffusion Models

2024-06-03 · CVPR 2024 1 · Qilong Zhangli, Jindong Jiang, Di Liu, Licheng Yu 외

While diffusion models have significantly advanced the quality of image generation their capability to accurately and coherently render text within these images remains a substantial challenge. Conventional diffusion-bas…

DiversityImage GenerationInstance SegmentationLayout Generation+2

Aligning where to see and what to tell: image caption with region-based attention and scene factorization

2015-06-20 · Junqi Jin, Kun fu, Runpeng Cui, Fei Sha 외

Recent progress on automatic generation of image captions has shown that it is possible to describe the most salient information conveyed by images with accurate and meaningful sentences. In this paper, we propose an ima…

Image Captioning

Word Searching in Scene Image and Video Frame in Multi-Script Scenario using Dynamic Shape Coding

2017-08-18 · Partha Pratim Roy, Ayan Kumar Bhunia, Avirup Bhattacharyya, Umapada Pal

Retrieval of text information from natural scene images and video frames is a challenging task due to its inherent problems like complex character shapes, low resolution, background noise, etc. Available OCR systems ofte…

Keyword SpottingOptical Character Recognition (OCR)RetrievalText Retrieval