paper-with-me

Papers

Optical Character Recognition using Convolutional Neural Networks for Ashokan Brahmi Inscriptions

2024-12-29 · Yash Agrawal, Srinidhi Balasubramanian, Rahul Meena, Rohail Alam, Himanshu Malviya, Rohini P

This research paper delves into the development of an Optical Character Recognition (OCR) system for the recognition of Ashokan Brahmi characters using Convolutional Neural Networks. It utilizes a comprehensive dataset of character images to train the models, along with data augmentation techniques to optimize the training process. Furthermore, the paper incorporates image preprocessing to remove noise, as well as image segmentation to facilitate line and character segmentation. The study mainly focuses on three pre-trained CNNs, namely LeNet, VGG-16, and MobileNet and compares their accuracy. Transfer learning was employed to adapt the pre-trained models to the Ashokan Brahmi character dataset. The findings reveal that MobileNet outperforms the other two models in terms of accuracy, achieving a validation accuracy of 95.94% and validation loss of 0.129. The paper provides an in-depth analysis of the implementation process using MobileNet and discusses the implications of the findings. The use of OCR for character recognition is of significant importance in the field of epigraphy, specifically for the preservation and digitization of ancient scripts. The results of this research paper demonstrate the effectiveness of using pre-trained CNNs for the recognition of Ashokan Brahmi characters.

📄 PDF Abstract BibTeX arXiv:2501.01981

Code (0)

등록된 구현이 없습니다.

Tasks

Data AugmentationImage SegmentationOptical Character RecognitionOptical Character Recognition (OCR)Semantic SegmentationTransfer Learning

Methods 이 논문이 사용한 방법론

VGG-16 설명 없음

Similar Papers 제목 키워드 기반

For the Purpose of Curry: A UD Treebank for Ashokan Prakrit

2021-11-24 · UDW (SyntaxFest) 2021 12 · Adam Farris, Aryaman Arora

We present the first linguistically annotated treebank of Ashokan Prakrit, an early Middle Indo-Aryan dialect continuum attested through Emperor Ashoka Maurya's 3rd century BCE rock and pillar edicts. For annotation, we …

Extensions to Brahmic script processing within the Nisaba library: new scripts, languages and utilities

2022-06-01 · LREC 2022 6 · Alexander Gutkin, Cibu Johny, Raiomond Doctor, Lawrence Wolf-Sonkin 외

The Brahmic family of scripts is used to record some of the most spoken languages in the world and is arguably the most diverse family of writing systems. In this work, we present several substantial extensions to Brahmi…

Transliteration

BrahmicTokenizer-131K: An Indic-Capable Drop-In Replacement for o200k_base

2026-05-28 · Rohan Shravan arxiv

We present BrahmicTokenizer-131K, a 131,072-vocabulary byte-level BPE tokenizer that closes the Brahmic compression gap at the 131K-vocabulary class while preserving the English, EU-language, and code compression of Open…

Finite-state script normalization and processing utilities: The Nisaba Brahmic library

2021-04-01 · EACL 2021 2 · Cibu Johny, Lawrence Wolf-Sonkin, Alexander Gutkin, Brian Roark

This paper presents an open-source library for efficient low-level processing of ten major South Asian Brahmic scripts. The library provides a flexible and extensible framework for supporting crucial operations on Brahmi…

SurveyTransliteration

Logios : An open source Greek Polytonic Optical Character Recognition system

2025-06-26 · Perifanos Konstantinos, Goutsos Dionisis

In this paper, we present an Optical Character Recognition (OCR) system specifically designed for the accurate recognition and digitization of Greek polytonic texts. By leveraging the combined strengths of convolutional …

Optical Character RecognitionOptical Character Recognition (OCR)