paper-with-me

Papers

Scene Text Magnifier

2019-06-17 · Toshiki Nakamura, Anna Zhu, Seiichi Uchida

Scene text magnifier aims to magnify text in natural scene images without recognition. It could help the special groups, who have myopia or dyslexia to better understand the scene. In this paper, we design the scene text magnifier through interacted four CNN-based networks: character erasing, character extraction, character magnify, and image synthesis. The architecture of the networks are extended based on the hourglass encoder-decoders. It inputs the original scene text image and outputs the text magnified image while keeps the background unchange. Intermediately, we can get the side-output results of text erasing and text extraction. The four sub-networks are first trained independently and fine-tuned in end-to-end mode. The training samples for each stage are processed through a flow with original image and text annotation in ICDAR2013 and Flickr dataset as input, and corresponding text erased image, magnified text annotation, and text magnified scene image as output. To evaluate the performance of text magnifier, the Structural Similarity is used to measure the regional changes in each character region. The experimental results demonstrate our method can magnify scene text effectively without effecting the background.

📄 PDF Abstract BibTeX arXiv:1907.00693

Code (0)

등록된 구현이 없습니다.

Tasks

Image Generationtext annotation

Similar Papers 제목 키워드 기반

Magnifier: A Multi-grained Neural Network-based Architecture for Burned Area Delineation

2025-04-28 · Daniele Rege Cambrin, Luca Colomba, Paolo Garza

In crisis management and remote sensing, image segmentation plays a crucial role, enabling tasks like disaster response and emergency planning by analyzing visual data. Neural networks are able to analyze satellite acqui…

Burned Area DelineationDisaster ResponseImage SegmentationSemantic Segmentation

LSceneLLM: Enhancing Large 3D Scene Understanding Using Adaptive Visual Preferences

2024-12-02 · CVPR 2025 1 · Hongyan Zhi, Peihao Chen, Junyan Li, Shuailei Ma 외

Research on 3D Vision-Language Models (3D-VLMs) is gaining increasing attention, which is crucial for developing embodied AI within 3D scenes, such as visual navigation and embodied question answering. Due to the high de…

Embodied Question AnsweringQuestion AnsweringScene UnderstandingVisual Navigation

Magnifier Prompt: Tackling Multimodal Hallucination via Extremely Simple Instructions

2024-10-15 · Yuhan Fu, Ruobing Xie, Jiazhen Liu, Bangxiang Lan 외

Hallucinations in multimodal large language models (MLLMs) hinder their practical applications. To address this, we propose a Magnifier Prompt (MagPrompt), a simple yet effective method to tackle hallucinations in MLLMs …

Hallucination

AdaZoom: Adaptive Zoom Network for Multi-Scale Object Detection in Large Scenes

2021-06-19 · Jingtao Xu, YaLi Li, Shengjin Wang

Detection in large-scale scenes is a challenging problem due to small objects and extreme scale variation. It is essential to focus on the image regions of small objects. In this paper, we propose a novel Adaptive Zoom (…

object-detectionObject Detection

MagnifierNet: Towards Semantic Adversary and Fusion for Person Re-identification

2020-02-25 · Yushi Lan, Yu-An Liu, Maoqing Tian, Xinchi Zhou 외

Although person re-identification (ReID) has achieved significant improvement recently by enforcing part alignment, it is still a challenging task when it comes to distinguishing visually similar identities or identifyin…

DiversityPerson Re-Identification