paper-with-me

홈 › Papers

Homograph Disambiguation Through Selective Diacritic Restoration

2019-12-10 · WS 2019 8 · Sawsan Alqahtani, Hanan Aldarmaki, Mona Diab

Lexical ambiguity, a challenging phenomenon in all natural languages, is particularly prevalent for languages with diacritics that tend to be omitted in writing, such as Arabic. Omitting diacritics leads to an increase in the number of homographs: different words with the same spelling. Diacritic restoration could theoretically help disambiguate these words, but in practice, the increase in overall sparsity leads to performance degradation in NLP applications. In this paper, we propose approaches for automatically marking a subset of words for diacritic restoration, which leads to selective homograph disambiguation. Compared to full or no diacritic restoration, these approaches yield selectively-diacritized datasets that balance sparsity and lexical disambiguation. We evaluate the various selection strategies extrinsically on several downstream applications: neural machine translation, part-of-speech tagging, and semantic textual similarity. Our experiments on Arabic show promising results, where our devised strategies on selective diacritization lead to a more balanced and consistent performance in downstream applications.

📄 PDF Abstract BibTeX arXiv:1912.04479

Code (0)

등록된 구현이 없습니다.

Tasks

Machine TranslationPart-Of-Speech TaggingSemantic Textual SimilarityTranslation

Similar Papers 제목 키워드 기반

A Novel Challenge Set for Hebrew Morphological Disambiguation and Diacritics Restoration

2020-10-06 · Findings of the Association for Computational Linguistics 2020 · Avi Shmidman, Joshua Guedalia, Shaltiel Shmidman, Moshe Koppel 외

One of the primary tasks of morphological parsers is the disambiguation of homographs. Particularly difficult are cases of unbalanced ambiguity, where one of the possible analyses is far more frequent than the others. In…

Morphological Disambiguation

Lexical Disambiguation of Igbo using Diacritic Restoration

2017-04-01 · WS 2017 4 · Ignatius Ezeani, Mark Hepple, Ikechukwu Onyenwe

Properly written texts in Igbo, a low-resource African language, are rich in both orthographic and tonal diacritics. Diacritics are essential in capturing the distinctions in pronunciation and meaning of words, as well a…

BIG-bench Machine LearningGeneral Classification

Attentive Sequence-to-Sequence Learning for Diacritic Restoration of Yorùbá Language Text

2018-04-03 · Iroro Orife

Yor\`ub\'a is a widely spoken West African language with a writing system rich in tonal and orthographic diacritics. With very few exceptions, diacritics are omitted from electronic texts, due to limited device and appli…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Machine Translationspeech-recognition+4

Correcting diacritics and typos with a ByT5 transformer model

2022-01-31 · Lukas Stankevičius, Mantas Lukoševičius, Jurgita Kapočiūtė-Dzikienė, Monika Briedienė 외

Due to the fast pace of life and online communications and the prevalence of English and the QWERTY keyboard, people tend to forgo using diacritics, make typographical errors (typos) when typing in other languages. Resto…

Improving Yorùbá Diacritic Restoration

2020-03-23 · Iroro Orife, David I. Adelani, Timi Fasubaa, Victor Williamson 외

Yor\`ub\'a is a widely spoken West African language with a writing system rich in orthographic and tonal diacritics. They provide morphological information, are crucial for lexical disambiguation, pronunciation and are v…