paper-with-me

홈 › Papers

Reconstructing Syllable Sequences in Abugida Scripts with Incomplete Inputs

2025-05-16 · Ye Kyaw Thu, Thazin Myint Oo

This paper explores syllable sequence prediction in Abugida languages using Transformer-based models, focusing on six languages: Bengali, Hindi, Khmer, Lao, Myanmar, and Thai, from the Asian Language Treebank (ALT) dataset. We investigate the reconstruction of complete syllable sequences from various incomplete input types, including consonant sequences, vowel sequences, partial syllables (with random character deletions), and masked syllables (with fixed syllable deletions). Our experiments reveal that consonant sequences play a critical role in accurate syllable prediction, achieving high BLEU scores, while vowel sequences present a significantly greater challenge. The model demonstrates robust performance across tasks, particularly in handling partial and masked syllable reconstruction, with strong results for tasks involving consonant information and syllable masking. This study advances the understanding of sequence prediction for Abugida languages and provides practical insights for applications such as text prediction, spelling correction, and data augmentation in these scripts.

📄 PDF Abstract BibTeX arXiv:2505.11008

Code (0)

등록된 구현이 없습니다.

Tasks

Data AugmentationPredictionSpelling Correction

Similar Papers 제목 키워드 기반

Orthographic Syllable as basic unit for SMT between Related Languages

2016-10-03 · EMNLP 2016 11 · Anoop Kunchukuttan, Pushpak Bhattacharyya

We explore the use of the orthographic syllable, a variable-length consonant-vowel sequence, as a basic unit of translation between related languages which use abugida or alphabetic scripts. We show that orthographic syl…

Translation

Unicode Normalization and Grapheme Parsing of Indic Languages

2023-05-11 · Nazmuddoha Ansary, Quazi Adibur Rahman Adib, Tahsin Reasat, Asif Shahriyar Sushmit 외

Writing systems of Indic languages have orthographic syllables, also known as complex graphemes, as unique horizontal units. A prominent feature of these languages is these complex grapheme units that comprise consonants…

Language Modelling

Separate Before You Compress: The WWHO Tokenization Architecture

2026-03-26 · Kusal Darshana arxiv

Current Large Language Models (LLMs) mostly use BPE (Byte Pair Encoding) based tokenizers, which are very effective for simple structured Latin scripts such as English. However, standard BPE tokenizers struggle to proces…

edATLAS: An Efficient Disambiguation Algorithm for Texting in Languages with Abugida Scripts

2021-01-05 · Sourav Ghosh, Sourabh Vasant Gothe, Chandramouli Sanchi, Barath Raj Kandur Raja

Abugida refers to a phonogram writing system where each syllable is represented using a single consonant or typographic ligature, along with a default vowel or optional diacritic(s) to denote other vowels. However, texti…

Language Modelling

Simplified Abugidas

2018-07-01 · ACL 2018 7 · Chenchen Ding, Masao Utiyama, Eiichiro Sumita

An abugida is a writing system where the consonant letters represent syllables with a default vowel and other vowels are denoted by diacritics. We investigate the feasibility of recovering the original text written in an…

Sentence