paper-with-me

Papers

The NaijaVoices Dataset: Cultivating Large-Scale, High-Quality, Culturally-Rich Speech Data for African Languages

2025-05-26 · Chris Emezue, The NaijaVoices Community, Busayo Awobade, Abraham Owodunni, Handel Emezue, Gloria Monica Tobechukwu Emezue, Nefertiti Nneoma Emezue, Sewade Ogun, Bunmi Akinremi, David Ifeoluwa Adelani, Chris Pal

The development of high-performing, robust, and reliable speech technologies depends on large, high-quality datasets. However, African languages -- including our focus, Igbo, Hausa, and Yoruba -- remain under-represented due to insufficient data. Popular voice-enabled technologies do not support any of the 2000+ African languages, limiting accessibility for circa one billion people. While previous dataset efforts exist for the target languages, they lack the scale and diversity needed for robust speech models. To bridge this gap, we introduce the NaijaVoices dataset, a 1,800-hour speech-text dataset with 5,000+ speakers. We outline our unique data collection approach, analyze its acoustic diversity, and demonstrate its impact through finetuning experiments on automatic speech recognition, averagely achieving 75.86% (Whisper), 52.06% (MMS), and 42.33% (XLSR) WER improvements. These results highlight NaijaVoices' potential to advance multilingual speech processing for African languages.

📄 PDF Abstract BibTeX arXiv:2505.20564

Code (0)

등록된 구현이 없습니다.

Tasks

Automatic Speech RecognitionDiversityspeech-recognitionSpeech Recognition

Similar Papers 제목 키워드 기반

A field guide to cultivating computational biology

2021-04-23 · Anne E Carpenter, Casey S Greene, Piero Carnici, Benilton S Carvalho 외

Biomedical research centers can empower basic discovery and novel therapeutic strategies by leveraging their large-scale datasets from experiments and patients. This data, together with new technologies to create and ana…

Cultivating DNN Diversity for Large Scale Video Labelling

2017-07-13 · Mikel Bober-Irizar, Sameed Husain, Eng-Jon Ong, Miroslaw Bober

We investigate factors controlling DNN diversity in the context of the Google Cloud and YouTube-8M Video Understanding Challenge. While it is well-known that ensemble methods improve prediction performance, and that comb…

DiversityVideo Understanding

Generating Translation Corpora in Indic Languages:Cultivating Bilingual Texts for Cross Lingual Fertilization

2015-12-01 · WS 2015 12 · Niladri Sekhar Dash, Arulmozi Selvraj, Mazhar Hussain
Translation

K-Myriad: Jump-starting reinforcement learning with unsupervised parallel agents

2026-01-26 · Vincenzo De Paola, Mirco Mutti, Riccardo Zamboni, Marcello Restelli arxiv

Parallelization in Reinforcement Learning is typically employed to speed up the training of a single policy, where multiple workers collect experience from an identical sampling distribution. This common design limits th…

Reinforcement LearningContinuous Control

VisReason: A Large-Scale Dataset for Visual Chain-of-Thought Reasoning

2025-11-21 · Lingxiao Li, Yifan Wang, Xinyan Gao, Chen Tang 외 arxiv

Chain-of-Thought (CoT) prompting has proven remarkably effective for eliciting complex reasoning in large language models (LLMs). Yet, its potential in multimodal large language models (MLLMs) remains largely untapped, h…

Visual Reasoning