paper-with-me

홈 › Papers

SERENGETI: Massively Multilingual Language Models for Africa

2022-12-21 · Ife Adebara, AbdelRahim Elmadany, Muhammad Abdul-Mageed, Alcides Alcoba Inciarte

Multilingual pretrained language models (mPLMs) acquire valuable, generalizable linguistic information during pretraining and have advanced the state of the art on task-specific finetuning. To date, only ~31 out of ~2,000 African languages are covered in existing language models. We ameliorate this limitation by developing SERENGETI, a massively multilingual language model that covers 517 African languages and language varieties. We evaluate our novel models on eight natural language understanding tasks across 20 datasets, comparing to 4 mPLMs that cover 4-23 African languages. SERENGETI outperforms other models on 11 datasets across the eights tasks, achieving 82.27 average F_1. We also perform analyses of errors from our models, which allows us to investigate the influence of language genealogy and linguistic similarity when the models are applied under zero-shot settings. We will publicly release our models for research.\footnote{\href{https://github.com/UBC-NLP/serengeti}{https://github.com/UBC-NLP/serengeti}}

📄 PDF Abstract BibTeX arXiv:2212.10785

Code (0)

등록된 구현이 없습니다.

Tasks

Language ModelingLanguage ModellingNatural Language Understanding

Similar Papers 제목 키워드 기반

From N-grams to Pre-trained Multilingual Models For Language Identification

2024-10-11 · Thapelo Sindane, Vukosi Marivate

In this paper, we investigate the use of N-gram models and Large Pre-trained Multilingual models for Language Identification (LID) across 11 South African languages. For N-gram models, this study shows that effective dat…

Language IdentificationXLM-R

Preparing the Vuk'uzenzele and ZA-gov-multilingual South African multilingual corpora

2023-03-07 · Richard Lastrucci, Isheanesu Dzingirai, Jenalea Rajab, Andani Madodonga 외

This paper introduces two multilingual government themed corpora in various South African languages. The corpora were collected by gathering the South African Government newspaper (Vuk'uzenzele), as well as South African…

Language ModelingLanguage ModellingMachine TranslationNMT+1

AfriHG: News headline generation for African Languages

2024-12-28 · Toyib Ogunremi, Serah Akojenu, Anthony Soronnadi, Olubayo Adekanmbi 외

This paper introduces AfriHG -- a news headline generation dataset created by combining from XLSum and MasakhaNEWS datasets focusing on 16 languages widely spoken by Africa. We experimented with two seq2eq models (mT5-ba…

Headline Generation

Designing and Contextualising Probes for African Languages

2025-05-15 · Wisdom Aduah, Francois Meyer

Pretrained language models (PLMs) for African languages are continually improving, but the reasons behind these advances remain unclear. This paper presents the first systematic investigation into probing PLMs for lingui…

Active LearningSentence

Cheetah: Natural Language Generation for 517 African Languages

2024-01-02 · Ife Adebara, AbdelRahim Elmadany, Muhammad Abdul-Mageed

Low-resource African languages pose unique challenges for natural language processing (NLP) tasks, including natural language generation (NLG). In this paper, we develop Cheetah, a massively multilingual NLG language mod…

DiversityLanguage ModelingLanguage ModellingText Generation