paper-with-me

홈 › Papers

Text Normalization for Low-Resource Languages of Africa

2021-03-29 · Andrew Zupon, Evan Crew, Sandy Ritchie

Training data for machine learning models can come from many different sources, which can be of dubious quality. For resource-rich languages like English, there is a lot of data available, so we can afford to throw out the dubious data. For low-resource languages where there is much less data available, we can't necessarily afford to throw out the dubious data, in case we end up with a training set which is too small to train a model. In this study, we examine the effects of text normalization and data set quality for a set of low-resource languages of Africa -- Afrikaans, Amharic, Hausa, Igbo, Malagasy, Somali, Swahili, and Zulu. We describe our text normalizer which we built in the Pynini framework, a Python library for finite state transducers, and our experiments in training language models for African languages using the Natural Language Toolkit (NLTK), an open-source Python library for NLP.

📄 PDF Abstract BibTeX arXiv:2103.15845

Code (0)

등록된 구현이 없습니다.

Tasks

Text Normalization

Similar Papers 제목 키워드 기반

A Sentiment Corpus for South African Under-Resourced Languages in a Multilingual Context

2022-06-01 · SIGUL (LREC) 2022 6 · Ronny Mabokela, Tim Schlippe

Multilingual sentiment analysis is a process of detecting and classifying sentiment based on textual information written in multiple languages. There has been tremendous research advancement on high-resourced languages s…

Sentiment Analysis

AFRICAPTION: Establishing a New Paradigm for Image Captioning in African Languages

2025-10-20 · Mardiyyah Oduwole, Prince Mireku, Fatimo Adebanjo, Oluwatosin Olajide 외 arxiv

Multimodal AI research has overwhelmingly focused on high-resource languages, hindering the democratization of advancements in the field. To address this, we present AfriCaption, a comprehensive framework for multilingua…

Image Captioning

Towards a parallel corpus of Portuguese and the Bantu language Emakhuwa of Mozambique

2021-04-12 · Felermino D. M. A. Ali, Andrew Caines, Jaimito L. A. Malavi

Major advancement in the performance of machine translation models has been made possible in part thanks to the availability of large-scale parallel corpora. But for most languages in the world, the existence of such cor…

Machine TranslationSentenceTranslation

Cheetah: Natural Language Generation for 517 African Languages

2024-01-02 · Ife Adebara, AbdelRahim Elmadany, Muhammad Abdul-Mageed

Low-resource African languages pose unique challenges for natural language processing (NLP) tasks, including natural language generation (NLG). In this paper, we develop Cheetah, a massively multilingual NLG language mod…

DiversityLanguage ModelingLanguage ModellingText Generation

Building Collaboration-based Resources in Endowed African Languages: Case of NTeALan Dictionaries Platform

2020-05-01 · LREC 2020 5 · Elvis Mboning Tchiaze, Jean Marc Bassahak, Daniel Baleba, W 외

In a context where open-source NLP resources and tools in African languages are scarce and dispersed, it is difficult for researchers to truly fit African languages into current algorithms of artificial intelligence. Cre…

Management