paper-with-me

Papers

Graphemic Normalization of the Perso-Arabic Script

2022-10-21 · Raiomond Doctor, Alexander Gutkin, Cibu Johny, Brian Roark, Richard Sproat

Since its original appearance in 1991, the Perso-Arabic script representation in Unicode has grown from 169 to over 440 atomic isolated characters spread over several code pages representing standard letters, various diacritics and punctuation for the original Arabic and numerous other regional orthographic traditions. This paper documents the challenges that Perso-Arabic presents beyond the best-documented languages, such as Arabic and Persian, building on earlier work by the expert community. We particularly focus on the situation in natural language processing (NLP), which is affected by multiple, often neglected, issues such as the use of visually ambiguous yet canonically nonequivalent letters and the mixing of letters from different orthographies. Among the contributing conflating factors are the lack of input methods, the instability of modern orthographies, insufficient literacy, and loss or lack of orthographic tradition. We evaluate the effects of script normalization on eight languages from diverse language families in the Perso-Arabic script diaspora on machine translation and statistical language modeling tasks. Our results indicate statistically significant improvements in performance in most conditions for all the languages considered when normalization is applied. We argue that better understanding and representation of Perso-Arabic script variation within regional orthographic traditions, where those are present, is crucial for further progress of modern computational NLP techniques especially for languages with a paucity of resources.

📄 PDF Abstract BibTeX arXiv:2210.12273

Code (1)

google-research/google-research/tree/master/perso_arabic_norm 공식 구현 jax

Tasks

Language ModelingLanguage ModellingMachine Translation

Similar Papers 제목 키워드 기반

Graphemic ambiguous queries on Arabic-scripted historical corpora

2019-09-01 · RANLP 2019 9 · Alicia Gonz{\'a}lez Mart{\'\i}nez

Beyond Arabic: Software for Perso-Arabic Script Manipulation

2023-01-26 · Alexander Gutkin, Cibu Johny, Raiomond Doctor, Brian Roark 외

This paper presents an open-source software library that provides a set of finite-state transducer (FST) components and corresponding utilities for manipulating the writing systems of languages that use the Perso-Arabic …

Transliteration

Automatic Long Audio Alignment and Confidence Scoring for Conversational Arabic Speech

2014-05-01 · LREC 2014 5 · Mohamed Elmahdy, Mark Hasegawa-Johnson, Eiman Mustafawi

In this paper, a framework for long audio alignment for conversational Arabic speech is proposed. Accurate alignments help in many speech processing tasks such as audio indexing, speech recognizer acoustic model (AM) tra…

Language Modellingspeech-recognitionSpeech Recognition

Script Normalization for Unconventional Writing of Under-Resourced Languages in Bilingual Communities

2023-05-25 · Sina Ahmadi, Antonios Anastasopoulos

The wide accessibility of social media has provided linguistically under-represented communities with an extraordinary opportunity to create content in their native languages. This, however, comes with certain challenges…

Language IdentificationMachine Translation

Phonetic and Graphemic Systems for Multi-Genre Broadcast Transcription

2018-02-01 · Yu Wang, Xie Chen, Mark Gales, Anton Ragni 외

State-of-the-art English automatic speech recognition systems typically use phonetic rather than graphemic lexicons. Graphemic systems are known to perform less well for English as the mapping from the written form to th…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)speech-recognitionSpeech Recognition