paper-with-me

Papers

BOTS-LM: Training Large Language Models for Setswana

2024-08-05 · Nathan Brown, Vukosi Marivate

In this work we present BOTS-LM, a series of bilingual language models proficient in both Setswana and English. Leveraging recent advancements in data availability and efficient fine-tuning, BOTS-LM achieves performance similar to models significantly larger than itself while maintaining computational efficiency. Our initial release features an 8 billion parameter generative large language model, with upcoming 0.5 billion and 1 billion parameter large language models and a 278 million parameter encoder-only model soon to be released. We find the 8 billion parameter model significantly outperforms Llama-3-70B and Aya 23 on English-Setswana translation tasks, approaching the performance of dedicated machine translation models, while approaching 70B parameter performance on Setswana reasoning as measured by a machine translated subset of the MMLU benchmark. To accompany the BOTS-LM series of language models, we release the largest Setswana web dataset, SetsText, totalling over 267 million tokens. In addition, we release the largest machine translated Setswana dataset, the first and largest synthetic Setswana dataset, training and evaluation code, training logs, and MMLU-tsn, a machine translated subset of MMLU.

📄 PDF Abstract BibTeX arXiv:2408.02239

Code (0)

등록된 구현이 없습니다.

Tasks

Computational EfficiencyLanguage ModelingLanguage ModellingLarge Language ModelMachine TranslationMMLUTranslation

Similar Papers 제목 키워드 기반

PuoBERTa: Training and evaluation of a curated language model for Setswana

2023-10-13 · Vukosi Marivate, Moseli Mots'oehli, Valencia Wagner, Richard Lastrucci 외

Natural language processing (NLP) has made significant progress for well-resourced languages such as English but lagged behind for low-resource languages like Setswana. This paper addresses this gap by presenting PuoBERT…

Language ModelingLanguage Modellingnamed-entity-recognitionNamed Entity Recognition+5

Training Cross-Lingual embeddings for Setswana and Sepedi

2021-11-11 · Mack Makgatho, Vukosi Marivate, Tshephisho Sefara, Valencia Wagner

African languages still lag in the advances of Natural Language Processing techniques, one reason being the lack of representative data, having a technique that can transfer information between languages can help mitigat…

Cross-Lingual TransferSemantic SimilaritySemantic Textual SimilarityWord Embeddings

Practical Approach on Implementation of WordNets for South African Languages

2021-01-01 · EACL (GWC) 2021 1 · Tshephisho Joseph Sefara, Tumisho Billson Mokgonyane, Vukosi Marivate

This paper proposes the implementation of WordNets for five South African languages, namely, Sepedi, Setswana, Tshivenda, isiZulu and isiXhosa to be added to open multilingual WordNets (OMW) on natural language toolkit (…

Language-Independent Sentiment Labelling with Distant Supervision: A Case Study for English, Sepedi and Setswana

2025-11-25 · Koena Ronny Mabokela, Tim Schlippe, Mpho Raborife, Turgay Celik arxiv

Sentiment analysis is a helpful task to automatically analyse opinions and emotions on various topics in areas such as AI for Social Good, AI in Education or marketing. While many of the sentiment analysis systems are de…

Sentiment Analysis

Investigating an approach for low resource language dataset creation, curation and classification: Setswana and Sepedi

2020-02-18 · LREC 2020 5 · Vukosi Marivate, Tshephisho Sefara, Vongani Chabalala, Keamogetswe Makhaya 외

The recent advances in Natural Language Processing have been a boon for well-represented languages in terms of available curated data and research resources. One of the challenges for low-resourced languages is clear gui…

ClassificationData AugmentationGeneral ClassificationTopic Classification