paper-with-me

홈 › Papers

Comparative Analysis of N-gram Text Representation on Igbo Text Document Similarity

2020-04-01 · Nkechi Ifeanyi-Reuben, Chidiebere Ugwu, Nwachukwu E. O

The improvement in Information Technology has encouraged the use of Igbo in the creation of text such as resources and news articles online. Text similarity is of great importance in any text-based applications. This paper presents a comparative analysis of n-gram text representation on Igbo text document similarity. It adopted Euclidean similarity measure to determine the similarities between Igbo text documents represented with two word-based n-gram text representation (unigram and bigram) models. The evaluation of the similarity measure is based on the adopted text representation models. The model is designed with Object-Oriented Methodology and implemented with Python programming language with tools from Natural Language Toolkits (NLTK). The result shows that unigram represented text has highest distance values whereas bigram has the lowest corresponding distance values. The lower the distance value, the more similar the two documents and better the quality of the model when used for a task that requires similarity measure. The similarity of two documents increases as the distance value moves down to zero (0). Ideally, the result analyzed revealed that Igbo text document similarity measured on bigram represented text gives accurate similarity result. This will give better, effective and accurate result when used for tasks such as text classification, clustering and ranking on Igbo text.

📄 PDF Abstract BibTeX arXiv:2004.00375

Code (0)

등록된 구현이 없습니다.

Tasks

ArticlesClusteringGeneral Classificationtext-classificationText Classificationtext similarity

Similar Papers 제목 키워드 기반

Analysis and representation of Igbo text document for a text-based system

2020-09-05 · Ifeanyi-Reuben Nkechi J., Ugwu Chidiebere, Adegbola Tunde

The advancement in Information Technology (IT) has assisted in inculcating the three Nigeria major languages in text-based application such as text mining, information retrieval and natural language processing. The inter…

Information RetrievalRetrieval

IgboBERT Models: Building and Training Transformer Models for the Igbo Language

2022-06-01 · LREC 2022 6 · Chiamaka Chukwuneke, Ignatius Ezeani, Paul Rayson, Mahmoud El-Haj

This work presents a standard Igbo named entity recognition (IgboNER) dataset as well as the results from training and fine-tuning state-of-the-art transformer IgboNER models. We discuss the process of our dataset creati…

Language ModelingLanguage Modellingnamed-entity-recognitionNamed Entity Recognition+1

Igbo Diacritic Restoration using Embedding Models

2018-06-01 · NAACL 2018 6 · Ignatius Ezeani, Mark Hepple, Ikechukwu Onyenwe, Enemouh Chioma

Igbo is a low-resource language spoken by approximately 30 million people worldwide. It is the native language of the Igbo people of south-eastern Nigeria. In Igbo language, diacritics - orthographic and tonal - play a h…

Machine TranslationWord Embeddings

The IgboAPI Dataset: Empowering Igbo Language Technologies through Multi-dialectal Enrichment

2024-05-02 · Chris Chinenye Emezue, Ifeoma Okoh, Chinedu Mbonu, Chiamaka Chukwuneke 외

The Igbo language is facing a risk of becoming endangered, as indicated by a 2025 UNESCO study. This highlights the need to develop language technologies for Igbo to foster communication, learning and preservation. To cr…

Machine TranslationTranslation

Development of a General Purpose Sentiment Lexicon for Igbo Language

2020-04-24 · WS 2019 8 · Emeka Ogbuju, Moses Onyesolu

There are publicly available general purpose sentiment lexicons in some high resource languages but very few exist in the low resource languages. This makes it difficult to directly perform sentiment analysis tasks in su…

Sentiment Analysis