paper-with-me

홈 › Papers

The Thiomi Dataset: A Large-Scale Multimodal Corpus for Low-Resource African Languages

2026-03-31 · Hillary Mutisya, John Mugane, Gavin Nyamboga, Brian Chege, Maryruth Gathoni arxiv

We present the Thiomi Dataset, a large-scale multimodal corpus spanning ten African languages across four language families: Swahili, Kikuyu, Kamba, Kimeru, Luo, Maasai, Kipsigis, Somali (East Africa); Wolof (West Africa); and Fulani (West/Central Africa). The dataset contains over 601,000 approved sentence-level text annotations and over 385,000 audio recordings, collected through a dedicated community data collection platform involving over 100 contributors. To validate the dataset's utility, we train and evaluate ASR, MT, and TTS models, establishing baselines across all languages. Our best ASR system achieves 3.24% WER on Swahili (Common Voice), reducing prior academic SOTA from 8.3% to 3.24% (5.1 percentage point absolute, 61% relative reduction), and 4.3% WER on Somali. The dataset will be published on HuggingFace. We describe the collection platform, quality assurance workflows, and baseline experiments, and discuss implications for African language technology infrastructure.

📄 PDF Abstract BibTeX arXiv:2603.29244

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

OmniCorpus: A Unified Multimodal Corpus of 10 Billion-Level Images Interleaved with Text

2024-06-12 · Qingyun Li, Zhe Chen, Weiyun Wang, Wenhai Wang 외

Image-text interleaved data, consisting of multiple images and texts arranged in a natural document format, aligns with the presentation paradigm of internet data and closely resembles human reading habits. Recent studie…

In-Context Learning

EgMM-Corpus: A Multimodal Vision-Language Dataset for Egyptian Culture

2025-10-17 · Mohamed Gamil, Abdelrahman Elsayed, Abdelrahman Lila, Ahmed Gad 외 arxiv

Despite recent advances in AI, multimodal culturally diverse datasets are still limited, particularly for regions in the Middle East and Africa. In this paper, we introduce EgMM-Corpus, a multimodal dataset dedicated to …

OceanPile: A Large-Scale Multimodal Ocean Corpus for Foundation Models

2026-04-25 · Yida Xue, Ningyu Zhang, Tingwei Wu, Zhe Ma 외 arxiv

The vast and underexplored ocean plays a critical role in regulating global climate and supporting marine biodiversity, yet artificial intelligence has so far delivered limited impact in this domain due to a fundamental …

A Large Scale Speech Sentiment Corpus

2020-05-01 · LREC 2020 5 · Eric Chen, Zhiyun Lu, Hao Xu, Liangliang Cao 외

We present a multimodal corpus for sentiment analysis based on the existing Switchboard-1 Telephone Speech Corpus released by the Linguistic Data Consortium. This corpus extends the Switchboard-1 Telephone Speech Corpus …

Sentiment Analysis

Multimodal Knowledge Learning for Named Entity Disambiguation

2021-08-17 · ACL ARR August 2021 8 · Anonymous

With the popularity of online social medias in recent years, massive-scale multimodal information has brought new challenges to traditional Named Entity Disambiguation (NED) tasks. Recently, Multimodal Named Entity Disam…

Entity DisambiguationMeta-LearningTransfer Learning