paper-with-me

홈 › Papers

NaijaNLP: A Survey of Nigerian Low-Resource Languages

2025-02-27 · Isa Inuwa-Dutse

With over 500 languages in Nigeria, three languages -- Hausa, Yor\ub\'a and Igbo -- spoken by over 175 million people, account for about 60% of the spoken languages. However, these languages are categorised as low-resource due to insufficient resources to support tasks in computational linguistics. Several research efforts and initiatives have been presented, however, a coherent understanding of the state of Natural Language Processing (NLP) - from grammatical formalisation to linguistic resources that support complex tasks such as language understanding and generation is lacking. This study presents the first comprehensive review of advancements in low-resource NLP (LR-NLP) research across the three major Nigerian languages (NaijaNLP). We quantitatively assess the available linguistic resources and identify key challenges. Although a growing body of literature addresses various NLP downstream tasks in Hausa, Igbo, and Yor\ub\'a, only about 25.1% of the reviewed studies contribute new linguistic resources. This finding highlights a persistent reliance on repurposing existing data rather than generating novel, high-quality resources. Additionally, language-specific challenges, such as the accurate representation of diacritics, remain under-explored. To advance NaijaNLP and LR-NLP more broadly, we emphasise the need for intensified efforts in resource enrichment, comprehensive annotation, and the development of open collaborative initiatives.

📄 PDF Abstract BibTeX arXiv:2502.19784

Code (0)

등록된 구현이 없습니다.

Tasks

Survey

Similar Papers 제목 키워드 기반

A Survey of Machine Translation Tasks on Nigerian Languages

2022-06-01 · LREC 2022 6 · Ebelechukwu Nwafor, Anietie Andy

Machine translation is an active area of research that has received a significant amount of attention over the past decade. With the advent of deep learning models, the translation of several languages has been performed…

Machine TranslationSurveyTranslation

Semi-automatic discourse annotation in a low-resource language: Developing a connective lexicon for Nigerian Pidgin

2021-11-01 · CODI 2021 11 · Marian Marchal, Merel Scholman, Vera Demberg

Cross-linguistic research on discourse structure and coherence marking requires discourse-annotated corpora and connective lexicons in a large number of languages. However, the availability of such resources is limited, …

Relation

NaijaNER : Comprehensive Named Entity Recognition for 5 Nigerian Languages

2021-03-30 · Wuraola Fisayo Oyewusi, Olubayo Adekanmbi, Ifeoma Okoh, Vitus Onuigwe 외

Most of the common applications of Named Entity Recognition (NER) is on English and other highly available languages. In this work, we present our findings on Named Entity Recognition for 5 Nigerian Languages (Nigerian E…

named-entity-recognitionNamed Entity RecognitionNamed Entity Recognition (NER)NER

Sometin Beta Pass Notin (SBPN): Improving Multilingual ASR for Nigerian Languages via Knowledge Distillation

2026-05-18 · Sewade Ogun arxiv

Although modern multilingual Automatic Speech Recognition (ASR) systems support several Nigerian languages, their performance consistently lags behind high-resource languages like English and French. Nigerian languages p…

Knowledge DistillationSpeech Recognition

NaijaSenti: A Nigerian Twitter Sentiment Corpus for Multilingual Sentiment Analysis

2022-01-20 · LREC 2022 6 · Shamsuddeen Hassan Muhammad, David Ifeoluwa Adelani, Sebastian Ruder, Ibrahim Said Ahmad 외

Sentiment analysis is one of the most widely studied applications in NLP, but most work focuses on languages with large amounts of data. We introduce the first large-scale human-annotated Twitter sentiment dataset for th…

Sentiment Analysis