A Multi-faceted OCR Framework for Artificial Urdu News Ticker Text Recognition
Content based information search and retrieval has allowed for easier access to data. While Latin based scripts have gained attention and support from academia and industry, there is limited support for cursive script languages, like Urdu. In this paper, we present the first instance of Urdu news ticker detection and recognition and take a micron sized step towards the goal of super intelligence. The presented solution allows for automating the transcription, indexing and captioning of Urdu news video content. We present the first comprehensive data set, to our knowledge, for Urdu news ticker recognition, collected from 41 different news channels. The data set covers both high and low quality channels, distorted and blurred news tickers, making the data set an ideal test case for any automatic Urdu News Recognition system in future. We identify and address the key challenges in Urdu News Ticker text recognition …
Code (0)
등록된 구현이 없습니다.
Tasks
Optical Character Recognition (OCR)RetrievalSimilar Papers 제목 키워드 기반
Clustering Urdu News Using Headlines
This paper that proposes and evaluates a new algorithm to automatically cluster Urdu news from different news agencies. The task is challenging because there are no language processing libraries for the Urdu language. Th…
ClusteringInformation RetrievalText ClusteringAx-to-Grind Urdu: Benchmark Dataset for Urdu Fake News Detection
Misinformation can seriously impact society, affecting anything from public opinion to institutional confidence and the political horizon of a state. Fake News (FN) proliferation on online websites and Online Social Netw…
Fact CheckingFake News DetectionMisinformationExploiting Transliterated Words for Finding Similarity in Inter-Language News Articles using Machine Learning
Finding similarities between two inter-language news articles is a challenging problem of Natural Language Processing (NLP). It is difficult to find similar news articles in a different language other than the native lan…
ArticlesMachine Translationtext-to-speechText to Speech+1Fake News Classification in Urdu: A Domain Adaptation Approach for a Low-Resource Language
Misinformation on social media is a widely acknowledged issue, and researchers worldwide are actively engaged in its detection. However, low-resource languages such as Urdu have received limited attention in this domain.…
News ClassificationDomain AdaptationUrdu News Article Recommendation Model using Natural Language Processing Techniques
There are several online newspapers in urdu but for the users it is difficult to find the content they are looking for because these most of them contain irrelevant data and most users did not get what they want to retri…
ArticlesLanguage ModelingLanguage Modelling