paper-with-me

Papers

Large Scale Scene Text Verification with Guided Attention

2018-04-23 · Dafang He, Yeqing Li, Alexander Gorban, Derrall Heath, Julian Ibarz, Qian Yu, Daniel Kifer, C. Lee Giles

Many tasks are related to determining if a particular text string exists in an image. In this work, we propose a new framework that learns this task in an end-to-end way. The framework takes an image and a text string as input and then outputs the probability of the text string being present in the image. This is the first end-to-end framework that learns such relationships between text and images in scene text area. The framework does not require explicit scene text detection or recognition and thus no bounding box annotations are needed for it. It is also the first work in scene text area that tackles suh a weakly labeled problem. Based on this framework, we developed a model called Guided Attention. Our designed model achieves much better results than several state-of-the-art scene text reading based solutions for a challenging Street View Business Matching task. The task tries to find correct business names for storefront images and the dataset we collected for it is substantially larger, and more challenging than existing scene text dataset. This new real-world task provides a new perspective for studying scene text related problems. We also demonstrate the uniqueness of our task via a comparison between our problem and a typical Visual Question Answering problem.

📄 PDF Abstract BibTeX arXiv:1804.08588

Code (0)

등록된 구현이 없습니다.

Tasks

Question AnsweringScene Text DetectionText DetectionVisual Question AnsweringVisual Question Answering (VQA)

Similar Papers 제목 키워드 기반

UniGeo: A Multi-modal Large Language Model for Text-Guided Cross-View Geo-Localization

2026-08-27 · Jiahao Wen, Hang Yu, Zhedong Zheng arxiv

Text-guided drone geo-localization aims to identify a target region in a large-scale image gallery from a natural-language description. Existing methods mainly formulate this task as direct matching between an open-ended…

Diagnosing Aerial-View Object Detectors with Foundational Image Generative Models

2026-07-02 · Stanislav Panev, Minhyek Jeon, Vaishnavi Khindkar, Ahish Deshpande 외 arxiv

Recent advances in large-scale image generative models enable photorealistic scene synthesis with controllable attributes. Beyond data augmentation, their potential as diagnostic tools for trained vision systems remains …

Data Augmentation

VeREFINE: Integrating Object Pose Verification with Physics-guided Iterative Refinement

2019-09-12 · Dominik Bauer, Timothy Patten, Markus Vincze

Accurate and robust object pose estimation for robotics applications requires verification and refinement steps. In this work, we propose to integrate hypotheses verification with object pose refinement guided by physics…

ObjectPose Estimation

GALA3D: Towards Text-to-3D Complex Scene Generation via Layout-guided Generative Gaussian Splatting

2024-02-11 · Xiaoyu Zhou, Xingjian Ran, Yajiao Xiong, Jinlin He 외

We present GALA3D, generative 3D GAussians with LAyout-guided control, for effective compositional text-to-3D generation. We first utilize large language models (LLMs) to generate the initial layout and introduce a layou…

3D GenerationScene GenerationText to 3D

A Training-Free Regeneration Paradigm: Contrastive Reflection Memory Guided Self-Verification and Self-Improvement

2026-03-20 · Yuran Li, Di Wu, Benoit Boulet arxiv

Verification-guided self-improvement has recently emerged as a promising approach to improving the accuracy of large language model (LLM) outputs. However, existing approaches face a trade-off between inference efficienc…