paper-with-me

Papers

GraDeT-HTR: A Resource-Efficient Bengali Handwritten Text Recognition System utilizing Grapheme-based Tokenizer and Decoder-only Transformer

2025-09-22 · Md. Mahmudul Hasan, Ahmed Nesar Tahsin Choudhury, Mahmudul Hasan, Md. Mosaddek Khan arxiv

Despite Bengali being the sixth most spoken language in the world, handwritten text recognition (HTR) systems for Bengali remain severely underdeveloped. The complexity of Bengali script--featuring conjuncts, diacritics, and highly variable handwriting styles--combined with a scarcity of annotated datasets makes this task particularly challenging. We present GraDeT-HTR, a resource-efficient Bengali handwritten text recognition system based on a Grapheme-aware Decoder-only Transformer architecture. To address the unique challenges of Bengali script, we augment the performance of a decoder-only transformer by integrating a grapheme-based tokenizer and demonstrate that it significantly improves recognition accuracy compared to conventional subword tokenizers. Our model is pretrained on large-scale synthetic data and fine-tuned on real human-annotated samples, achieving state-of-the-art performance on multiple benchmark datasets.

📄 PDF Abstract BibTeX arXiv:2509.18081

Code (0)

등록된 구현이 없습니다.

Tasks

Handwritten Text Recognition

Similar Papers 제목 키워드 기반

BeHGAN: Bengali Handwritten Word Generation from Plain Text Using Generative Adversarial Networks

2025-12-25 · Md. Rakibul Islam, Md. Kamrozzaman Bhuiyan, Safwan Muntasir, Arifur Rahman Jawad 외 arxiv

Handwritten Text Recognition (HTR) is a well-established research area. In contrast, Handwritten Text Generation (HTG) is an emerging field with significant potential. This task is challenging due to the variation in ind…

Handwritten Text RecognitionText Generation

Optical Text Recognition in Nepali and Bengali: A Transformer-based Approach

2024-04-03 · S M Rakib Hasan, Aakar Dhakal, Md Humaion Kabir Mehedi, Annajiat Alim Rasel

Efforts on the research and development of OCR systems for Low-Resource Languages are relatively new. Low-resource languages have little training data available for training Machine Translation systems or other systems. …

DecoderMachine TranslationOptical Character Recognition (OCR)

Bengali Handwritten Digit Recognition using CNN with Explainable AI

2022-12-23 · MD Tanvir Rouf Shawon, Raihan Tanvir, Md. Golam Rabiul Alam

Handwritten character recognition is a hot topic for research nowadays. If we can convert a handwritten piece of paper into a text-searchable document using the Optical Character Recognition (OCR) technique, we can easil…

Explainable Artificial Intelligence (XAI)Handwritten Digit RecognitionOptical Character RecognitionOptical Character Recognition (OCR)

End-to-End Optical Character Recognition for Bengali Handwritten Words

2021-05-09 · Farisa Benta Safir, Abu Quwsar Ohi, M. F. Mridha, Muhammad Mostafa Monowar 외

Optical character recognition (OCR) is a process of converting analogue documents into digital using document images. Currently, many commercial and non-commercial OCR systems exist for both handwritten and printed copie…

Optical Character RecognitionOptical Character Recognition (OCR)

Multichannel Attention Networks with Ensembled Transfer Learning to Recognize Bangla Handwritten Charecter

2024-08-20 · Farhanul Haque, Md. Al-Hasan, Sumaiya Tabssum Mou, Abu Saleh Musa Miah 외

The Bengali language is the 5th most spoken native and 7th most spoken language in the world, and Bengali handwritten character recognition has attracted researchers for decades. However, other languages such as English,…

Handwriting RecognitionTransfer Learning