paper-with-me

Papers

TextBlockV2: Towards Precise-Detection-Free Scene Text Spotting with Pre-trained Language Model

2024-03-15 · Jiahao Lyu, Jin Wei, Gangyan Zeng, Zeng Li, Enze Xie, Wei Wang, Yu Zhou

Existing scene text spotters are designed to locate and transcribe texts from images. However, it is challenging for a spotter to achieve precise detection and recognition of scene texts simultaneously. Inspired by the glimpse-focus spotting pipeline of human beings and impressive performances of Pre-trained Language Models (PLMs) on visual tasks, we ask: 1) "Can machines spot texts without precise detection just like human beings?", and if yes, 2) "Is text block another alternative for scene text spotting other than word or character?" To this end, our proposed scene text spotter leverages advanced PLMs to enhance performance without fine-grained detection. Specifically, we first use a simple detector for block-level text detection to obtain rough positional information. Then, we finetune a PLM using a large-scale OCR dataset to achieve accurate recognition. Benefiting from the comprehensive language knowledge gained during the pre-training phase, the PLM-based recognition module effectively handles complex scenarios, including multi-line, reversed, occluded, and incomplete-detection texts. Taking advantage of the fine-tuned language model on scene recognition benchmarks and the paradigm of text block detection, extensive experiments demonstrate the superior performance of our scene text spotter across multiple public benchmarks. Additionally, we attempt to spot texts directly from an entire scene image to demonstrate the potential of PLMs, even Large Language Models (LLMs).

📄 PDF Abstract BibTeX arXiv:2403.10047

Code (0)

등록된 구현이 없습니다.

Tasks

Language ModelingLanguage ModellingOptical Character Recognition (OCR)Scene RecognitionText DetectionText Spotting

Similar Papers 제목 키워드 기반

SynthText3D: Synthesizing Scene Text Images from 3D Virtual Worlds

2019-07-13 · Minghui Liao, Boyu Song, Shangbang Long, Minghang He 외

With the development of deep neural networks, the demand for a significant amount of annotated training data becomes the performance bottlenecks in many fields of research and applications. Image synthesis can generate a…

Image GenerationScene Text DetectionText Detection

The Scene Language: Representing Scenes with Programs, Words, and Embeddings

2024-10-22 · CVPR 2025 1 · Yunzhi Zhang, Zizhang Li, Matt Zhou, Shangzhe Wu 외

We introduce the Scene Language, a visual scene representation that concisely and precisely describes the structure, semantics, and identity of visual scenes. It represents a scene with three key components: a program th…

Scene Generation

SceneTextStylizer: A Training-Free Scene Text Style Transfer Framework with Diffusion Model

2025-10-13 · Honghui Yuan, Keiji Yanai arxiv

With the rapid development of diffusion models, style transfer has made remarkable progress. However, flexible and localized style editing for scene text remains an unsolved challenge. Although existing scene text editin…

Text Style Transfer

UnrealText: Synthesizing Realistic Scene Text Images from the Unreal World

2020-03-24 · CVPR 2020 6 · Shangbang Long, Cong Yao

Synthetic data has been a critical tool for training scene text detection and recognition models. On the one hand, synthetic word images have proven to be a successful substitute for real images in training scene text re…

Image GenerationScene Text DetectionScene Text RecognitionText Detection

SCLARO: A Dataset for Grounded Scenario-Level Scene Understanding and ScenarioCLIP for Benchmarking

2025-11-25 · Advik Sinha, Saurabh Atreya, Aashutosh A, Sk Aziz Ali 외 arxiv

In the paradigm of computer vision-based precise real-world scene understanding, joint reasoning in terms of contextual understanding about the objects present in a scene, their inter-object relations, and the action bei…

Knowledge DistillationGraph ClassificationScene UnderstandingObject Detection