paper-with-me

홈 › Papers

Koshur Diacritizer: A Byte-Level Sequence-to-Sequence Model for Kashmiri Diacritic Restoration

2026-06-14 · Haq Nawaz Malik, Nahfid Nissar, Faizan Iqbal arxiv

Kashmiri, an Indo-Aryan language written in a modified Perso-Arabic script, frequently omits diacritic marks in digital text, creating ambiguity and challenging downstream NLP applications. We present Koshur Diacritizer, a ByT5-small byte-level sequence-to-sequence model for restoring diacritics in Kashmiri text. To support this task, we release a publicly available dataset of 23.7k aligned undiacritized diacritized Kashmiri sentence pairs. The proposed framework combines script-aware normalization, alignment validation, and skeleton-preserving inference to ensure reliable restoration while maintaining the original base-letter sequence. Experimental results on a held-out test set achieve a DERm of 0.2012 and a WER of 0.2159. Additionally, evaluation by a native Kashmiri linguistic expert yields a mean accuracy of 77.5%. The dataset, model, and source code are publicly released to provide a reproducible baseline for Kashmiri diacritic restoration and future low-resource language research.

📄 PDF Abstract BibTeX arXiv:2606.15883

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Arabic Diacritization: Stats, Rules, and Hacks

2017-04-01 · WS 2017 4 · Kareem Darwish, Hamdy Mubarak, Ahmed Abdelali

In this paper, we present a new and fast state-of-the-art Arabic diacritizer that guesses the diacritics of words and then their case endings. We employ a Viterbi decoder at word-level with back-off to stem, morphologica…

DecoderPart-Of-Speech TaggingTransliterationWord Sense Disambiguation

MEGABYTE: Predicting Million-byte Sequences with Multiscale Transformers

2023-05-12 · NeurIPS 2023 11 · Lili Yu, Dániel Simig, Colin Flaherty, Armen Aghajanyan 외

Autoregressive transformers are spectacular models for short sequences but scale poorly to long sequences such as high-resolution images, podcasts, code, or books. We proposed Megabyte, a multi-scale decoder architecture…

DecoderDensity EstimationLanguage ModelingLanguage Modelling

ByT5: Towards a token-free future with pre-trained byte-to-byte models

2021-05-28 · Linting Xue, Aditya Barua, Noah Constant, Rami Al-Rfou 외

Most widely-used pre-trained language models operate on sequences of tokens corresponding to word or subword units. By comparison, token-free models that operate directly on raw text (bytes or characters) have many benef…

Cross-Lingual Natural Language InferenceCross-Lingual NERCross-Lingual Paraphrase IdentificationCross-Lingual Question Answering+2

Understanding the Natural Language of DNA using Encoder-Decoder Foundation Models with Byte-level Precision

2023-11-04 · Aditya Malusare, Harish Kothandaraman, Dipesh Tamboli, Nadia A. Lanman 외

This paper presents the Ensemble Nucleotide Byte-level Encoder-Decoder (ENBED) foundation model, analyzing DNA sequences at byte-level precision with an encoder-decoder Transformer architecture. ENBED uses a sub-quadrati…

DecoderLanguage ModelingLanguage ModellingMasked Language Modeling

byteSteady: Fast Classification Using Byte-Level n-Gram Embeddings

2021-06-24 · Xiang Zhang, Alexandre Drouin, Raymond Li

This article introduces byteSteady -- a fast model for classification using byte-level n-gram embeddings. byteSteady assumes that each input comes as a sequence of bytes. A representation vector is produced using the ave…

Classificationtext-classificationText Classification