paper-with-me

홈 › Papers

Illegible Text to Readable Text: An Image-to-Image Transformation using Conditional Sliced Wasserstein Adversarial Networks

2019-10-11 · Mostafa Karimi, Gopalkrishna Veni, Yen-Yun Yu

Automatic text recognition from ancient handwritten record images is an important problem in the genealogy domain. However, critical challenges such as varying noise conditions, vanishing texts, and variations in handwriting make the recognition task difficult. We tackle this problem by developing a handwritten-to-machine-print conditional Generative Adversarial network (HW2MP-GAN) model that formulates handwritten recognition as a text-Image-to-text-Image translation problem where a given image, typically in an illegible form, is converted into another image, close to its machine-print form. The proposed model consists of three-components including a generator, and word-level and character-level discriminators. The model incorporates Sliced Wasserstein distance (SWD) and U-Net architectures in HW2MP-GAN for better quality image-to-image transformation. Our experiments reveal that HW2MP-GAN outperforms state-of-the-art baseline cGAN models by almost 30 in Frechet Handwritten Distance (FHD), 0.6 on average Levenshtein distance and 39% in word accuracy for image-to-image translation on IAM database. Further, HW2MP-GAN improves handwritten recognition word accuracy by 1.3% compared to baseline handwritten recognition models on the IAM database.

📄 PDF Abstract BibTeX arXiv:1910.05425

Code (0)

등록된 구현이 없습니다.

Tasks

Generative Adversarial NetworkImage-to-Image TranslationImage to textTranslation

Methods 이 논문이 사용한 방법론

Concatenated Skip Connection A Concatenated Skip Connection is a type of skip connection that seeks to reuse features by concatenating them to new layers, allowing more information to be retained from…
ReLU How Do I Communicate to Expedia? How Do I Communicate to Expedia? – Call ☎️ +1-(888) 829 (0881) or +1-805-330-4056 or +1-805-330-4056 for Live Support & Special Travel…
Max Pooling Max Pooling is a pooling operation that calculates the maximum value for patches of a feature map, and uses it to create a downsampled (pooled) feature map. It is usually…
Convolution A convolution is a type of matrix operation, consisting of a kernel, a small matrix of weights, that slides over input data performing element-wise multiplication with the…
U-Net 설명 없음

Similar Papers 제목 키워드 기반

Forensic License Plate Recognition with Compression-Informed Transformers

2022-07-29 · Denise Moussa, Anatol Maier, Andreas Spruck, Jürgen Seiler 외

Forensic license plate recognition (FLPR) remains an open challenge in legal contexts such as criminal investigations, where unreadable license plates (LPs) need to be deciphered from highly compressed and/or low resolut…

License Plate Recognition

Reasoning Models Sometimes Output Illegible Chains of Thought

2025-10-31 · Arun Jose arxiv

Language models trained via outcome-based reinforcement learning (RL) to reason using chain-of-thought (CoT) have shown remarkable performance. Monitoring such a model's CoT may allow us to understand its intentions and …

Reinforcement Learning

LPLC: A Dataset for License Plate Legibility Classification

2025-08-25 · Lucas Wojcik, Gabriel E. Lima, Valfride Nascimento, Eduil Nascimento 외 arxiv

Automatic License Plate Recognition (ALPR) faces a major challenge when dealing with illegible license plates (LPs). While reconstruction methods such as super-resolution (SR) have emerged, the core issue of recognizing …

License Plate RecognitionComputational Efficiency

COCO-Text: Dataset and Benchmark for Text Detection and Recognition in Natural Images

2016-01-26 · Andreas Veit, Tomas Matera, Lukas Neumann, Jiri Matas 외

This paper describes the COCO-Text dataset. In recent years large-scale datasets like SUN and Imagenet drove the advancement of scene understanding and object recognition. The goal of COCO-Text is to advance state-of-the…

DiversityGeneral ClassificationObject RecognitionOptical Character Recognition+4

Towards Image-based Automatic Meter Reading in Unconstrained Scenarios: A Robust and Efficient Approach

2020-09-21 · Rayson Laroca, Alessandra B. Araujo, Luiz A. Zanlorensi, Eduardo C. de Almeida 외

Existing approaches for image-based Automatic Meter Reading (AMR) have been evaluated on images captured in well-controlled scenarios. However, real-world meter reading presents unconstrained scenarios that are way more …

Image-based Automatic Meter ReadingMeter ReadingOptical Character Recognition (OCR)