paper-with-me

홈 › Papers

SwahBERT: Language Model of Swahili

2022-07-01 · NAACL 2022 7 · Gati Martin, Medard Edmund Mswahili, Young-Seob Jeong, Jeong Young-Seob

The rapid development of social networks, electronic commerce, mobile Internet, and other technologies, has influenced the growth of Web data.Social media and Internet forums are valuable sources of citizens’ opinions, which can be analyzed for community development and user behavior analysis.Unfortunately, the scarcity of resources (i.e., datasets or language models) become a barrier to the development of natural language processing applications in low-resource languages.Thanks to the recent growth of online forums and news platforms of Swahili, we introduce two datasets of Swahili in this paper: a pre-training dataset of approximately 105MB with 16M words and annotated dataset of 13K instances for the emotion classification task.The emotion classification dataset is manually annotated by two native Swahili speakers.We pre-trained a new monolingual language model for Swahili, namely SwahBERT, using our collected pre-training data, and tested it with four downstream tasks including emotion classification.We found that SwahBERT outperforms multilingual BERT, a well-known existing language model, in almost all downstream tasks.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Tasks

Emotion ClassificationLanguage ModelingLanguage Modellingmodel

Similar Papers 제목 키워드 기반

Not All Pretraining are Created Equal: Threshold Tuning and Class Weighting for Imbalanced Polarization Tasks in Low-Resource Settings

2026-03-08 · Abass Oguntade arxiv

This paper describes my submission to the Polarization Shared Task at SemEval-2025, which addresses polarization detection and classification in social media text. I develop Transformer-based systems for English and Swah…

The First Swahili Language Scene Text Detection and Recognition Dataset

2024-05-19 · Fadila Wendigoundi Douamba, Jianjun Song, Ling Fu, Yuliang Liu 외

Scene text recognition is essential in many applications, including automated translation, information retrieval, driving assistance, and enhancing accessibility for individuals with visual impairments. Much research has…

Information RetrievalScene Text DetectionScene Text RecognitionText Detection

Phonemic Representation and Transcription for Speech to Text Applications for Under-resourced Indigenous African Languages: The Case of Kiswahili

2022-10-29 · Ebbie Awino, Lilian Wanzare, Lawrence Muchemi, Barack Wanjawa 외

Building automatic speech recognition (ASR) systems is a challenging task, especially for under-resourced languages that need to construct corpora nearly from scratch and lack sufficient training data. It has emerged tha…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)speech-recognitionSpeech Recognition+1

KenSwQuAD -- A Question Answering Dataset for Swahili Low Resource Language

2022-05-04 · Barack W. Wanjawa, Lilian D. A. Wanzare, Florence Indede, Owen McOnyango 외

The need for Question Answering datasets in low resource languages is the motivation of this research, leading to the development of Kencorpus Swahili Question Answering Dataset, KenSwQuAD. This dataset is annotated from…

BIG-bench Machine LearningQuestion AnsweringReading Comprehension

SwaQuAD-24: QA Benchmark Dataset in Swahili

2024-10-18 · Alfred Malengo Kondoro

This paper proposes the creation of a Swahili Question Answering (QA) benchmark dataset, aimed at addressing the underrepresentation of Swahili in natural language processing (NLP). Drawing from established benchmarks li…

DiversityInformation RetrievalMachine TranslationQuestion Answering+1