paper-with-me

홈 › Papers

Unconstrained Scene Text and Video Text Recognition for Arabic Script

2017-11-07 · Mohit Jain, Minesh Mathew, C. V. Jawahar

Building robust recognizers for Arabic has always been challenging. We demonstrate the effectiveness of an end-to-end trainable CNN-RNN hybrid architecture in recognizing Arabic text in videos and natural scenes. We outperform previous state-of-the-art on two publicly available video text datasets - ALIF and ACTIV. For the scene text recognition task, we introduce a new Arabic scene text dataset and establish baseline results. For scripts like Arabic, a major challenge in developing robust recognizers is the lack of large quantity of annotated data. We overcome this by synthesising millions of Arabic text images from a large vocabulary of Arabic words and phrases. Our implementation is built on top of the model introduced here [37] which is proven quite effective for English scene text recognition. The model follows a segmentation-free, sequence to sequence transcription approach. The network transcribes a sequence of convolutional features from the input image to a sequence of target labels. This does away with the need for segmenting input image into constituent characters/glyphs, which is often difficult for Arabic script. Further, the ability of RNNs to model contextual dependencies yields superior recognition results.

📄 PDF Abstract BibTeX arXiv:1711.02396

Code (0)

등록된 구현이 없습니다.

Tasks

Scene Text Recognition

Similar Papers 제목 키워드 기반

RoadText-1K: Text Detection & Recognition Dataset for Driving Videos

2020-05-19 · Sangeeth Reddy, Minesh Mathew, Lluis Gomez, Marcal Rusinol 외

Perceiving text is crucial to understand semantics of outdoor scenes and hence is a critical requirement to build intelligent systems for driver assistance and self-driving. Most of the existing datasets for text detecti…

Text Detection

An Automatic System for Unconstrained Video-Based Face Recognition

2018-12-10 · Jingxiao Zheng, Rajeev Ranjan, Ching-Hui Chen, Jun-Cheng Chen 외

Although deep learning approaches have achieved performance surpassing humans for still image-based face recognition, unconstrained video-based face recognition is still a challenging task due to large volume of data to …

Face Recognition

Team LEYA in 10th ABAW Competition: Multimodal Ambivalence/Hesitancy Recognition Approach

2026-03-13 · Elena Ryumina, Alexandr Axyonov, Dmitry Sysoev, Timur Abdulkadirov 외 arxiv

Ambivalence/hesitancy recognition in unconstrained videos is a challenging problem due to the subtle, multimodal, and context-dependent nature of this behavioral state. In this paper, a multimodal approach for video-leve…

Sequence to sequence learning for unconstrained scene text recognition

2016-07-20 · Ahmed Mamdouh A. Hassanien

In this work we present a state-of-the-art approach for unconstrained natural scene text recognition. We propose a cascade approach that incorporates a convolutional neural network (CNN) architecture followed by a long s…

Scene Text Recognition

Accurate, Data-Efficient, Unconstrained Text Recognition with Convolutional Neural Networks

2018-12-31 · Mohamed Yousef, Khaled F. Hussain, Usama S. Mohammed

Unconstrained text recognition is an important computer vision task, featuring a wide variety of different sub-tasks, each with its own set of challenges. One of the biggest promises of deep neural networks has been the …

Handwriting RecognitionLicense Plate RecognitionOptical Character Recognition (OCR)Scene Text Recognition