paper-with-me

홈 › Papers

SwaQuAD-24: QA Benchmark Dataset in Swahili

2024-10-18 · Alfred Malengo Kondoro

This paper proposes the creation of a Swahili Question Answering (QA) benchmark dataset, aimed at addressing the underrepresentation of Swahili in natural language processing (NLP). Drawing from established benchmarks like SQuAD, GLUE, KenSwQuAD, and KLUE, the dataset will focus on providing high-quality, annotated question-answer pairs that capture the linguistic diversity and complexity of Swahili. The dataset is designed to support a variety of applications, including machine translation, information retrieval, and social services like healthcare chatbots. Ethical considerations, such as data privacy, bias mitigation, and inclusivity, are central to the dataset development. Additionally, the paper outlines future expansion plans to include domain-specific content, multimodal integration, and broader crowdsourcing efforts. The Swahili QA dataset aims to foster technological innovation in East Africa and provide an essential resource for NLP research and applications in low-resource languages.

📄 PDF Abstract BibTeX arXiv:2410.14289

Code (0)

등록된 구현이 없습니다.

Tasks

DiversityInformation RetrievalMachine TranslationQuestion AnsweringRetrieval

Methods 이 논문이 사용한 방법론

Focus 설명 없음

Similar Papers 제목 키워드 기반

The First Swahili Language Scene Text Detection and Recognition Dataset

2024-05-19 · Fadila Wendigoundi Douamba, Jianjun Song, Ling Fu, Yuliang Liu 외

Scene text recognition is essential in many applications, including automated translation, information retrieval, driving assistance, and enhancing accessibility for individuals with visual impairments. Much research has…

Information RetrievalScene Text DetectionScene Text RecognitionText Detection

Artificially Fluent: Swahili AI Performance Benchmarks Between English-Trained and Natively-Trained Datasets

2025-09-03 · Sophie Jaffer, Simeon Sayer arxiv

As large language models (LLMs) expand multilingual capabilities, questions remain about the equity of their performance across languages. While many communities stand to benefit from AI systems, the dominance of English…

SwahBERT: Language Model of Swahili

2022-07-01 · NAACL 2022 7 · Gati Martin, Medard Edmund Mswahili, Young-Seob Jeong, Jeong Young-Seob

The rapid development of social networks, electronic commerce, mobile Internet, and other technologies, has influenced the growth of Web data.Social media and Internet forums are valuable sources of citizens’ opinions, w…

Emotion ClassificationLanguage ModelingLanguage Modellingmodel

KenSwQuAD -- A Question Answering Dataset for Swahili Low Resource Language

2022-05-04 · Barack W. Wanjawa, Lilian D. A. Wanzare, Florence Indede, Owen McOnyango 외

The need for Question Answering datasets in low resource languages is the motivation of this research, leading to the development of Kencorpus Swahili Question Answering Dataset, KenSwQuAD. This dataset is annotated from…

BIG-bench Machine LearningQuestion AnsweringReading Comprehension

MMTAfrica: Multilingual Machine Translation for African Languages

2022-04-08 · WMT (EMNLP) 2021 11 · Chris C. Emezue, Bonaventure F. P. Dossou

In this paper, we focus on the task of multilingual machine translation for African languages and describe our contribution in the 2021 WMT Shared Task: Large-Scale Multilingual Machine Translation. We introduce MMTAfric…

Machine TranslationTranslation