paper-with-me

홈 › Papers

Corpus-Based Approaches to Igbo Diacritic Restoration

2026-01-26 · Ignatius Ezeani arxiv

With natural language processing (NLP), researchers aim to enable computers to identify and understand patterns in human languages. This is often difficult because a language embeds many dynamic and varied properties in its syntax, pragmatics and phonology, which need to be captured and processed. The capacity of computers to process natural languages is increasing because NLP researchers are pushing its boundaries. But these research works focus more on well-resourced languages such as English, Japanese, German, French, Russian, Mandarin Chinese, etc. Over 95% of the world's 7000 languages are low-resourced for NLP, i.e. they have little or no data, tools, and techniques for NLP work. In this thesis, we present an overview of diacritic ambiguity and a review of previous diacritic disambiguation approaches on other languages. Focusing on the Igbo language, we report the steps taken to develop a flexible framework for generating datasets for diacritic restoration. Three main approaches, the standard n-gram model, the classification models and the embedding models were proposed. The standard n-gram models use a sequence of previous words to the target stripped word as key predictors of the correct variants. For the classification models, a window of words on both sides of the target stripped word was used. The embedding models compare the similarity scores of the combined context word embeddings and the embeddings of each of the candidate variant vectors.

📄 PDF Abstract BibTeX arXiv:2601.18380

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Igbo Diacritic Restoration using Embedding Models

2018-06-01 · NAACL 2018 6 · Ignatius Ezeani, Mark Hepple, Ikechukwu Onyenwe, Enemouh Chioma

Igbo is a low-resource language spoken by approximately 30 million people worldwide. It is the native language of the Igbo people of south-eastern Nigeria. In Igbo language, diacritics - orthographic and tonal - play a h…

Machine TranslationWord Embeddings

Lexical Disambiguation of Igbo using Diacritic Restoration

2017-04-01 · WS 2017 4 · Ignatius Ezeani, Mark Hepple, Ikechukwu Onyenwe

Properly written texts in Igbo, a low-resource African language, are rich in both orthographic and tonal diacritics. Diacritics are essential in capturing the distinctions in pronunciation and meaning of words, as well a…

BIG-bench Machine LearningGeneral Classification

Transferred Embeddings for Igbo Similarity, Analogy, and Diacritic Restoration Tasks

2018-08-01 · COLING 2018 8 · Ignatius Ezeani, Ikechukwu Onyenwe, Mark Hepple

Existing NLP models are mostly trained with data from well-resourced languages. Most minority languages face the challenge of lack of resources - data and technologies - for NLP research. Building these resources from sc…

Transfer LearningWord EmbeddingsWord Similarity

Igbo-English Machine Translation: An Evaluation Benchmark

2020-04-01 · Ignatius Ezeani, Paul Rayson, Ikechukwu Onyenwe, Chinedu Uchechukwu 외

Although researchers and practitioners are pushing the boundaries and enhancing the capacities of NLP tools and methods, works on African languages are lagging. A lot of focus on well resourced languages such as English,…

Machine TranslationPart-Of-Speech TaggingTranslation

Corpus-Based Diacritic Restoration for South Slavic Languages

2016-05-01 · LREC 2016 5 · Nikola Ljube{\v{s}}i{\'c}, Toma{\v{z}} Erjavec, Darja Fi{\v{s}}er

In computer-mediated communication, Latin-based scripts users often omit diacritics when writing. Such text is typically easily understandable to humans but very difficult for computational processing because many words …