paper-with-me

홈 › Papers

Artificially Fluent: Swahili AI Performance Benchmarks Between English-Trained and Natively-Trained Datasets

2025-09-03 · Sophie Jaffer, Simeon Sayer arxiv

As large language models (LLMs) expand multilingual capabilities, questions remain about the equity of their performance across languages. While many communities stand to benefit from AI systems, the dominance of English in training data risks disadvantaging non-English speakers. To test the hypothesis that such data disparities may affect model performance, this study compares two monolingual BERT models: one trained and tested entirely on Swahili data, and another on comparable English news data. To simulate how multilingual LLMs process non-English queries through internal translation and abstraction, we translated the Swahili news data into English and evaluated it using the English-trained model. This approach tests the hypothesis by evaluating whether translating Swahili inputs for evaluation on an English model yields better or worse performance compared to training and testing a model entirely in Swahili, thus isolating the effect of language consistency versus cross-lingual abstraction. The results prove that, despite high-quality translation, the native Swahili-trained model performed better than the Swahili-to-English translated model, producing nearly four times fewer errors: 0.36% vs. 1.47% respectively. This gap suggests that translation alone does not bridge representational differences between languages and that models trained in one language may struggle to accurately interpret translated inputs due to imperfect internal knowledge representation, suggesting that native-language training remains important for reliable outcomes. In educational and informational contexts, even small performance gaps may compound inequality. Future research should focus on addressing broader dataset development for underrepresented languages and renewed attention to multilingual model evaluation, ensuring the reinforcing effect of global AI deployment on existing digital divides is reduced.

📄 PDF Abstract BibTeX arXiv:2509.04516

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

SwaQuAD-24: QA Benchmark Dataset in Swahili

2024-10-18 · Alfred Malengo Kondoro

This paper proposes the creation of a Swahili Question Answering (QA) benchmark dataset, aimed at addressing the underrepresentation of Swahili in natural language processing (NLP). Drawing from established benchmarks li…

DiversityInformation RetrievalMachine TranslationQuestion Answering+1

Graph Convolutional Network for Swahili News Classification

2021-03-16 · Alexandros Kastanos, Tyler Martin

This work empirically demonstrates the ability of Text Graph Convolutional Network (Text GCN) to outperform traditional natural language processing benchmarks for the task of semi-supervised Swahili news classification. …

ClassificationGeneral ClassificationNews Classification

MMTAfrica: Multilingual Machine Translation for African Languages

2022-04-08 · WMT (EMNLP) 2021 11 · Chris C. Emezue, Bonaventure F. P. Dossou

In this paper, we focus on the task of multilingual machine translation for African languages and describe our contribution in the 2021 WMT Shared Task: Large-Scale Multilingual Machine Translation. We introduce MMTAfric…

Machine TranslationTranslation

The First Swahili Language Scene Text Detection and Recognition Dataset

2024-05-19 · Fadila Wendigoundi Douamba, Jianjun Song, Ling Fu, Yuliang Liu 외

Scene text recognition is essential in many applications, including automated translation, information retrieval, driving assistance, and enhancing accessibility for individuals with visual impairments. Much research has…

Information RetrievalScene Text DetectionScene Text RecognitionText Detection

Phonemic Representation and Transcription for Speech to Text Applications for Under-resourced Indigenous African Languages: The Case of Kiswahili

2022-10-29 · Ebbie Awino, Lilian Wanzare, Lawrence Muchemi, Barack Wanjawa 외

Building automatic speech recognition (ASR) systems is a challenging task, especially for under-resourced languages that need to construct corpora nearly from scratch and lack sufficient training data. It has emerged tha…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)speech-recognitionSpeech Recognition+1