paper-with-me

Papers

VL-Reader: Vision and Language Reconstructor is an Effective Scene Text Recognizer

2024-09-18 · Humen Zhong, Zhibo Yang, Zhaohai Li, Peng Wang, Jun Tang, Wenqing Cheng, Cong Yao

Text recognition is an inherent integration of vision and language, encompassing the visual texture in stroke patterns and the semantic context among the character sequences. Towards advanced text recognition, there are three key challenges: (1) an encoder capable of representing the visual and semantic distributions; (2) a decoder that ensures the alignment between vision and semantics; and (3) consistency in the framework during pre-training, if it exists, and fine-tuning. Inspired by masked autoencoding, a successful pre-training strategy in both vision and language, we propose an innovative scene text recognition approach, named VL-Reader. The novelty of the VL-Reader lies in the pervasive interplay between vision and language throughout the entire process. Concretely, we first introduce a Masked Visual-Linguistic Reconstruction (MVLR) objective, which aims at simultaneously modeling visual and linguistic information. Then, we design a Masked Visual-Linguistic Decoder (MVLD) to further leverage masked vision-language context and achieve bi-modal feature interaction. The architecture of VL-Reader maintains consistency from pre-training to fine-tuning. In the pre-training stage, VL-Reader reconstructs both masked visual and text tokens, while in the fine-tuning stage, the network degrades to reconstruct all characters from an image without any masked regions. VL-reader achieves an average accuracy of 97.1% on six typical datasets, surpassing the SOTA by 1.1%. The improvement was even more significant on challenging datasets. The results demonstrate that vision and language reconstructor can serve as an effective scene text recognizer.

📄 PDF Abstract BibTeX arXiv:2409.11656

Code (0)

등록된 구현이 없습니다.

Tasks

DecoderScene Text Recognition

Similar Papers 제목 키워드 기반

Hierarchical Photo-Scene Encoder for Album Storytelling

2019-02-02 · Bairui Wang, Lin Ma, Wei zhang, Wenhao Jiang 외

In this paper, we propose a novel model with a hierarchical photo-scene encoder and a reconstructor for the task of album storytelling. The photo-scene encoder contains two sub-encoders, namely the photo and scene encode…

DecoderImage-guided Story Ending GenerationVisual Storytelling

AREA3D: Active Reconstruction Agent with Unified Feed-Forward 3D Perception and Vision-Language Guidance

2025-11-28 · Tianling Xu, Shengzhe Gan, Leslie Gu, Yuelei Li 외 arxiv

Active 3D reconstruction enables an agent to autonomously select viewpoints to efficiently obtain accurate and complete scene geometry, rather than passively reconstructing scenes from pre-collected images. However, exis…

3D Reconstruction

CLIP4STR: A Simple Baseline for Scene Text Recognition with Pre-trained Vision-Language Model

2023-05-23 · Shuai Zhao, Ruijie Quan, Linchao Zhu, Yi Yang

Pre-trained vision-language models~(VLMs) are the de-facto foundation models for various downstream tasks. However, scene text recognition methods still prefer backbones pre-trained on a single modality, namely, the visu…

DecoderLanguage ModelingLanguage ModellingScene Text Recognition

InSPECtor: an end-to-end design framework for compressive pixelated hyperspectral instruments

2023-09-19 · T. A. Stockmans, F. Snik, M. Esposito, C. van Dijk 외

Classic designs of hyperspectral instrumentation densely sample the spatial and spectral information of the scene of interest. Data may be compressed after the acquisition. In this paper we introduce a framework for the …

VGGT-$Ω$

2026-05-14 · Jianyuan Wang, Minghao Chen, Shangzhan Zhang, Nikita Karaev 외 arxiv

Recent feed-forward reconstruction models, such as VGGT, have proven competitive with traditional optimization-based reconstructors while also providing geometry-aware features useful for other tasks. Here, we show that …

Self-Supervised Learning