paper-with-me

Papers

Script Normalization for Unconventional Writing of Under-Resourced Languages in Bilingual Communities

2023-05-25 · Sina Ahmadi, Antonios Anastasopoulos

The wide accessibility of social media has provided linguistically under-represented communities with an extraordinary opportunity to create content in their native languages. This, however, comes with certain challenges in script normalization, particularly where the speakers of a language in a bilingual community rely on another script or orthography to write their native language. This paper addresses the problem of script normalization for several such languages that are mainly written in a Perso-Arabic script. Using synthetic data with various levels of noise and a transformer-based model, we demonstrate that the problem can be effectively remediated. We conduct a small-scale evaluation of real data as well. Our experiments indicate that script normalization is also beneficial to improve the performance of downstream tasks such as machine translation and language identification.

📄 PDF Abstract BibTeX arXiv:2305.16407

Code (1)

sinaahmadi/scriptnormalization 공식 구현

Tasks

Language IdentificationMachine Translation

Similar Papers 제목 키워드 기반

A Case Against Implicit Standards: Homophone Normalization in Machine Translation for Languages that use the Ge'ez Script

2025-07-20 · Hellina Hailu Nigatu, Atnafu Lambebo Tonja, Henok Biadglign Ademtew, Hizkel Mitiku Alemayehu 외 arxiv

Homophone normalization, where characters that have the same sound in a writing script are mapped to one character, is a pre-processing step applied in Amharic Natural Language Processing (NLP) literature. While this may…

Cross-Lingual TransferMachine TranslationTransfer Learning

PALI: A Language Identification Benchmark for Perso-Arabic Scripts

2023-04-03 · Sina Ahmadi, Milind Agarwal, Antonios Anastasopoulos

The Perso-Arabic scripts are a family of scripts that are widely adopted and used by various linguistic communities around the globe. Identifying various languages using such scripts is crucial to language technologies a…

Language Identification

Transformer based Urdu Handwritten Text Optical Character Reader

2022-06-09 · Mohammad Daniyal Shaiq, Musa Dildar Ahmed Cheema, Ali Kamal

Extracting Handwritten text is one of the most important components of digitizing information and making it available for large scale setting. Handwriting Optical Character Reader (OCR) is a research problem in computer …

Natural Language UnderstandingOptical Character Recognition (OCR)Position

GeezSwitch: Language Identification in Typologically Related Low-resourced East African Languages

2022-06-01 · LREC 2022 6 · Fitsum Gaim, Wonsuk Yang, Jong C. Park

Language identification is one of the fundamental tasks in natural language processing that is a prerequisite to data processing and numerous applications. Low-resourced languages with similar typologies are generally co…

Language IdentificationMachine Translation

A Clustering Framework for Lexical Normalization of Roman Urdu

2020-03-31 · Abdul Rafae Khan, Asim Karim, Hassan Sajjad, Faisal Kamiran 외

Roman Urdu is an informal form of the Urdu language written in Roman script, which is widely used in South Asia for online textual content. It lacks standard spelling and hence poses several normalization challenges duri…

ClusteringLexical Normalization