DiffusionSTR: Diffusion Model for Scene Text Recognition
This paper presents Diffusion Model for Scene Text Recognition (DiffusionSTR), an end-to-end text recognition framework using diffusion models for recognizing text in the wild. While existing studies have viewed the scene text recognition task as an image-to-text transformation, we rethought it as a text-text one under images in a diffusion model. We show for the first time that the diffusion model can be applied to text recognition. Furthermore, experimental results on publicly available datasets show that the proposed method achieves competitive accuracy compared to state-of-the-art methods.
Code (0)
등록된 구현이 없습니다.
Tasks
Image to textmodelScene Text RecognitionMethods 이 논문이 사용한 방법론
Similar Papers 제목 키워드 기반
Recognition-Guided Diffusion Model for Scene Text Image Super-Resolution
Scene Text Image Super-Resolution (STISR) aims to enhance the resolution and legibility of text within low-resolution (LR) images, consequently elevating recognition accuracy in Scene Text Recognition (STR). Previous met…
DenoisingDiversityImage Super-ResolutionScene Text Recognition+1On Manipulating Scene Text in the Wild with Diffusion Models
Diffusion models have gained attention for image editing yielding impressive results in text-to-image tasks. On the downside, one might notice that generated images of stable diffusion models suffer from deteriorated det…
Optical Character RecognitionOptical Character Recognition (OCR)Scene Text EditingTextSSR: Diffusion-based Data Synthesis for Scene Text Recognition
Scene text recognition (STR) suffers from the challenges of either less realistic synthetic training data or the difficulty of collecting sufficient high-quality real-world data, limiting the effectiveness of trained STR…
Image GenerationOptical Character Recognition (OCR)Scene Text EditingScene Text RecognitionDiffusion in the Dark: A Diffusion Model for Low-Light Text Recognition
Capturing images is a key part of automation for high-level tasks such as scene text recognition. Low-light conditions pose a challenge for high-level perception stacks, which are often optimized on well-lit, artifact-fr…
Image ReconstructionScene Text RecognitionLayout Agnostic Scene Text Image Synthesis with Diffusion Models
While diffusion models have significantly advanced the quality of image generation their capability to accurately and coherently render text within these images remains a substantial challenge. Conventional diffusion-bas…
DiversityImage GenerationInstance SegmentationLayout Generation+2