Optical Text Recognition in Nepali and Bengali: A Transformer-based Approach
Efforts on the research and development of OCR systems for Low-Resource Languages are relatively new. Low-resource languages have little training data available for training Machine Translation systems or other systems. Even though a vast amount of text has been digitized and made available on the internet the text is still in PDF and Image format, which are not instantly accessible. This paper discusses text recognition for two scripts: Bengali and Nepali; there are about 300 and 40 million Bengali and Nepali speakers respectively. In this study, using encoder-decoder transformers, a model was developed, and its efficacy was assessed using a collection of optical text images, both handwritten and printed. The results signify that the suggested technique corresponds with current approaches and achieves high precision in recognizing text in Bengali and Nepali. This study can pave the way for the advanced and accessible study of linguistics in South East Asia.
Code (0)
등록된 구현이 없습니다.
Tasks
DecoderMachine TranslationOptical Character Recognition (OCR)Similar Papers 제목 키워드 기반
GraDeT-HTR: A Resource-Efficient Bengali Handwritten Text Recognition System utilizing Grapheme-based Tokenizer and Decoder-only Transformer
Despite Bengali being the sixth most spoken language in the world, handwritten text recognition (HTR) systems for Bengali remain severely underdeveloped. The complexity of Bengali script--featuring conjuncts, diacritics,…
Handwritten Text RecognitionOptimizing Nepali PDF Extraction: A Comparative Study of Parser and OCR Technologies
This research compares PDF parsing and Optical Character Recognition (OCR) methods for extracting Nepali content from PDFs. PDF parsing offers fast and accurate extraction but faces challenges with non-Unicode Nepali fon…
Optical Character RecognitionOptical Character Recognition (OCR)End-to-End Optical Character Recognition for Bengali Handwritten Words
Optical character recognition (OCR) is a process of converting analogue documents into digital using document images. Currently, many commercial and non-commercial OCR systems exist for both handwritten and printed copie…
Optical Character RecognitionOptical Character Recognition (OCR)bbOCR: An Open-source Multi-domain OCR Pipeline for Bengali Documents
Despite the existence of numerous Optical Character Recognition (OCR) tools, the lack of comprehensive open-source systems hampers the progress of document digitization in various low-resource languages, including Bengal…
distortion correctionOptical Character RecognitionOptical Character Recognition (OCR)Bengali Handwritten Digit Recognition using CNN with Explainable AI
Handwritten character recognition is a hot topic for research nowadays. If we can convert a handwritten piece of paper into a text-searchable document using the Optical Character Recognition (OCR) technique, we can easil…
Explainable Artificial Intelligence (XAI)Handwritten Digit RecognitionOptical Character RecognitionOptical Character Recognition (OCR)