paper-with-me

Papers

A Large Multi-Target Dataset of Common Bengali Handwritten Graphemes

2020-10-01 · Samiul Alam, Tahsin Reasat, Asif Shahriyar Sushmit, Sadi Mohammad Siddiquee, Fuad Rahman, Mahady Hasan, Ahmed Imtiaz Humayun

Latin has historically led the state-of-the-art in handwritten optical character recognition (OCR) research. Adapting existing systems from Latin to alpha-syllabary languages is particularly challenging due to a sharp contrast between their orthographies. The segmentation of graphical constituents corresponding to characters becomes significantly hard due to a cursive writing system and frequent use of diacritics in the alpha-syllabary family of languages. We propose a labeling scheme based on graphemes (linguistic segments of word formation) that makes segmentation in-side alpha-syllabary words linear and present the first dataset of Bengali handwritten graphemes that are commonly used in an everyday context. The dataset contains 411k curated samples of 1295 unique commonly used Bengali graphemes. Additionally, the test set contains 900 uncommon Bengali graphemes for out of dictionary performance evaluation. The dataset is open-sourced as a part of a public Handwritten Grapheme Classification Challenge on Kaggle to benchmark vision algorithms for multi-target grapheme classification. The unique graphemes present in this dataset are selected based on commonality in the Google Bengali ASR corpus. From competition proceedings, we see that deep-learning methods can generalize to a large span of out of dictionary graphemes which are absent during training. Dataset and starter codes at www.kaggle.com/c/bengaliai-cv19.

📄 PDF Abstract BibTeX arXiv:2010.00170

Code (2)

BengaliAI/graphemePrepare 공식 구현
bhuiyanmobasshir94/Bengali.AI-Handwritten-Grapheme-Classification pytorch

Tasks

Multi-Label ClassificationOptical Character RecognitionOptical Character Recognition (OCR)

Similar Papers 제목 키워드 기반

Bengali Common Voice Speech Dataset for Automatic Speech Recognition

2022-06-28 · Samiul Alam, Asif Sushmit, Zaowad Abdullah, Shahrin Nakkhatra 외

Bengali is one of the most spoken languages in the world with over 300 million speakers globally. Despite its popularity, research into the development of Bengali speech recognition systems is hindered due to the lack of…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)DiversitySentence+2

Synthetic Error Dataset Generation Mimicking Bengali Writing Pattern

2020-03-07 · Md. Habibur Rahman Sifat, Chowdhury Rafeed Rahman, Mohammad Rafsan, Md. Hasibur Rahman

While writing Bengali using English keyboard, users often make spelling mistakes. The accuracy of any Bengali spell checker or paragraph correction module largely depends on the kind of error dataset it is based on. Manu…

Dataset Generation

Bangla-Wave: Improving Bangla Automatic Speech Recognition Utilizing N-gram Language Models

2022-09-13 · Mohammed Rakib, Md. Ismail Hossain, Nabeel Mohammed, Fuad Rahman

Although over 300M around the world speak Bangla, scant work has been done in improving Bangla voice-to-text transcription due to Bangla being a low-resource language. However, with the introduction of the Bengali Common…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Language ModelingLanguage Modelling+2

BengaliFig: A Low-Resource Challenge for Figurative and Culturally Grounded Reasoning in Bengali

2025-11-25 · Abdullah Al Sefat arxiv

Large language models excel on broad multilingual benchmarks but remain to be evaluated extensively in figurative and culturally grounded reasoning, especially in low-resource contexts. We present BengaliFig, a compact y…

Deciphering Hate: Identifying Hateful Memes and Their Targets

2024-03-16 · Eftekhar Hossain, Omar Sharif, Mohammed Moshiul Hoque, Sarah M. Preum

Internet memes have become a powerful means for individuals to express emotions, thoughts, and perspectives on social media. While often considered as a source of humor and entertainment, memes can also disseminate hatef…