paper-with-me

홈 › Papers

CLIPTER: Looking at the Bigger Picture in Scene Text Recognition

2023-01-18 · ICCV 2023 1 · Aviad Aberdam, David Bensaïd, Alona Golts, Roy Ganz, Oren Nuriel, Royee Tichauer, Shai Mazor, Ron Litman

Reading text in real-world scenarios often requires understanding the context surrounding it, especially when dealing with poor-quality text. However, current scene text recognizers are unaware of the bigger picture as they operate on cropped text images. In this study, we harness the representative capabilities of modern vision-language models, such as CLIP, to provide scene-level information to the crop-based recognizer. We achieve this by fusing a rich representation of the entire image, obtained from the vision-language model, with the recognizer word-level features via a gated cross-attention mechanism. This component gradually shifts to the context-enhanced representation, allowing for stable fine-tuning of a pretrained recognizer. We demonstrate the effectiveness of our model-agnostic framework, CLIPTER (CLIP TExt Recognition), on leading text recognition architectures and achieve state-of-the-art results across multiple benchmarks. Furthermore, our analysis highlights improved robustness to out-of-vocabulary words and enhanced generalization in low-data regimes.

📄 PDF Abstract BibTeX arXiv:2301.07464

Code (0)

등록된 구현이 없습니다.

Tasks

Language ModelingLanguage ModellingScene Text Recognition

Methods 이 논문이 사용한 방법론

CLIP Contrastive Language-Image Pre-training (CLIP), consisting of a simplified version of ConVIRT trained from scratch, is an efficient method of image representation learning…

Similar Papers 제목 키워드 기반

Seeing the Bigger Picture: 3D Latent Mapping for Mobile Manipulation Policy Learning

2025-10-04 · Sunghwan Kim, Woojeh Chung, Zhirui Dai, Dwait Bhatt 외 arxiv

In this paper, we demonstrate that mobile manipulation policies utilizing a 3D latent map achieve stronger spatial and temporal reasoning than policies relying solely on images. We introduce Seeing the Bigger Picture (SB…

Reinforcement Learning

More Words and Bigger Pictures

2013-06-01 · SEMEVAL 2013 6 · David Forsyth
Object Recognition

Reduction of Class Activation Uncertainty with Background Information

2023-05-05 · H M Dipu Kabir

Multitask learning is a popular approach to training high-performing neural networks with improved generalization. In this paper, we propose a background class to achieve improved generalization at a lower computation co…

ClassificationFine-Grained Image ClassificationImage ClassificationSatellite Image Classification

Looking Outside the Window: Wide-Context Transformer for the Semantic Segmentation of High-Resolution Remote Sensing Images

2021-06-29 · Lei Ding, Dong Lin, Shaofu Lin, Jing Zhang 외

Long-range contextual information is crucial for the semantic segmentation of High-Resolution (HR) Remote Sensing Images (RSIs). However, image cropping operations, commonly used for training neural networks, limit the p…

Image CroppingSemantic Segmentation

DREAM: Improving Situational QA by First Elaborating the Situation

2021-12-16 · NAACL 2022 7 · Yuling Gu, Bhavana Dalvi Mishra, Peter Clark

When people answer questions about a specific situation, e.g., "I cheated on my mid-term exam last week. Was that wrong?", cognitive science suggests that they form a mental picture of that situation before answering. Wh…

Question Answering