paper-with-me

Papers

How Good are Commercial Large Language Models on African Languages?

2023-05-11 · Jessica Ojo, Kelechi Ogueji

Recent advancements in Natural Language Processing (NLP) has led to the proliferation of large pretrained language models. These models have been shown to yield good performance, using in-context learning, even on unseen tasks and languages. They have also been exposed as commercial APIs as a form of language-model-as-a-service, with great adoption. However, their performance on African languages is largely unknown. We present a preliminary analysis of commercial large language models on two tasks (machine translation and text classification) across eight African languages, spanning different language families and geographical areas. Our results suggest that commercial language models produce below-par performance on African languages. We also find that they perform better on text classification than machine translation. In general, our findings present a call-to-action to ensure African languages are well represented in commercial large language models, given their growing popularity.

📄 PDF Abstract BibTeX arXiv:2305.06530

Code (0)

등록된 구현이 없습니다.

Tasks

In-Context LearningLanguage ModelingLanguage ModellingMachine Translationtext-classificationText ClassificationTranslation

Similar Papers 제목 키워드 기반

How good are Large Language Models on African Languages?

2023-11-14 · Jessica Ojo, Kelechi Ogueji, Pontus Stenetorp, David Ifeoluwa Adelani

Recent advancements in natural language processing have led to the proliferation of large language models (LLMs). These models have been shown to yield good performance, using in-context learning, even on tasks and langu…

In-Context LearningLanguage ModellingMachine Translationnamed-entity-recognition+6

The African Language Tax: Quantifying the Cost, Latency, and Context Penalty of Tokenizing African Languages in Frontier LLMs

2026-06-23 · Olaoye Anthony Somide arxiv

Commercial large language models bill, scale latency, and budget context per token. Yet tokenizers assign more subword tokens to the same meaning in some languages than in others, so speakers of languages with high token…

DN at SemEval-2023 Task 12: Low-Resource Language Text Classification via Multilingual Pretrained Language Model Fine-tuning

2023-05-04 · Daniil Homskiy, Narek Maloyan

In recent years, sentiment analysis has gained significant importance in natural language processing. However, most existing models and datasets for sentiment analysis are developed for high-resource languages, such as E…

Language ModelingLanguage ModellingSentiment Analysistext-classification+2

Building Collaboration-based Resources in Endowed African Languages: Case of NTeALan Dictionaries Platform

2020-05-01 · LREC 2020 5 · Elvis Mboning Tchiaze, Jean Marc Bassahak, Daniel Baleba, W 외

In a context where open-source NLP resources and tools in African languages are scarce and dispersed, it is difficult for researchers to truly fit African languages into current algorithms of artificial intelligence. Cre…

Management

AfroDigits: A Community-Driven Spoken Digit Dataset for African Languages

2023-03-22 · Chris Chinenye Emezue, Sanchit Gandhi, Lewis Tunstall, Abubakar Abid 외

The advancement of speech technologies has been remarkable, yet its integration with African languages remains limited due to the scarcity of African speech corpora. To address this issue, we present AfroDigits, a minima…