Improving Amharic Handwritten Word Recognition Using Auxiliary Task
Amharic is one of the official languages of the Federal Democratic Republic of Ethiopia. It is one of the languages that use an Ethiopic script which is derived from Gee'z, ancient and currently a liturgical language. Amharic is also one of the most widely used literature-rich languages of Ethiopia. There are very limited innovative and customized research works in Amharic optical character recognition (OCR) in general and Amharic handwritten text recognition in particular. In this study, Amharic handwritten word recognition will be investigated. State-of-the-art deep learning techniques including convolutional neural networks together with recurrent neural networks and connectionist temporal classification (CTC) loss were used to make the recognition in an end-to-end fashion. More importantly, an innovative way of complementing the loss function using the auxiliary task from the row-wise similarities of the Amharic alphabet was tested to show a significant recognition improvement over a baseline method. Such findings will promote innovative problem-specific solutions as well as will open insight to a generalized solution that emerges from problem-specific domains.
Code (0)
등록된 구현이 없습니다.
Tasks
Handwritten Text RecognitionOptical Character RecognitionOptical Character Recognition (OCR)Similar Papers 제목 키워드 기반
Handwritten Amharic Character Recognition Using a Convolutional Neural Network
Amharic is the official language of the Federal Democratic Republic of Ethiopia. There are lots of historic Amharic and Ethiopic handwritten documents addressing various relevant issues including governance, science, rel…
Data AugmentationMulti-Task LearningHandwritten Amharic Character Recognition System Using Convolutional Neural Networks
Amharic language is an official language of the federal government of the Federal Democratic Republic of Ethiopia. Accordingly, there is a bulk of handwritten Amharic documents available in libraries, information centres…
BIG-bench Machine LearningRetrievalOffline Handwritten Amharic Character Recognition Using Few-shot Learning
Few-shot learning is an important, but challenging problem of machine learning aimed at learning from only fewer labeled training examples. It has become an active area of research due to deep learning requiring huge amo…
Few-Shot Learningimage-classificationImage ClassificationAmharic-English Speech Translation in Tourism Domain
This paper describes speech translation from Amharic-to-English, particularly Automatic Speech Recognition (ASR) with post-editing feature and Amharic-English Statistical Machine Translation (SMT). ASR experiment is cond…
Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Language ModelingLanguage Modelling+4Multi-script Handwritten Digit Recognition Using Multi-task Learning
Handwritten digit recognition is one of the extensively studied area in machine learning. Apart from the wider research on handwritten digit recognition on MNIST dataset, there are many other research works on various sc…
Handwritten Digit RecognitionMulti-Task Learning