paper-with-me

홈 › Papers

From Speech to Text Corpora: Evaluating ASR-Based Data Acquisition for Low-Resource Fongbe and Hausa

2026-06-20 · Mahounan Pericles Adjovi, Victor Olufemi, Roald Eiselen, Prasenjit Mitra arxiv

Low-resource African languages lack text corpora needed for language model training. We investigate whether ASR pipelines can extend text resources for two typologically distinct West African languages: Fongbe (tonal, diacritic-rich) and Hausa (non-tonal). We fine-tune MMS-300M on a curated 12.3-hour Fongbe dataset, achieving 9.48% WER on the ALFFA benchmark - a 78% relative reduction from the prior 44.04% baseline - while preserving tonal diacritics critical to the language. For Hausa, we apply an existing fine-tuned Whisper-Small model. We catalog 1,553 YouTube videos (236 hours) and process a subset of 424 videos (45.49 hours) selected to balance domain diversity with available computational resources, producing 6,770 transcribed segments. Human evaluation on 50 randomly sampled segments per language shows mean quality scores of 57.4/100 for Hausa and 36.5/100 for Fongbe, indicating that while Hausa transcriptions approach acceptable quality for corpus construction, Fongbe transcriptions require post-processing or improved models for production use. We release the curated dataset, fine-tuned model, transcribed corpus, and full video catalog following platform terms and ethical guidelines.

📄 PDF Abstract BibTeX arXiv:2606.22274

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Context-based Word Acquisition for Situated Dialogue in a Virtual World

2014-01-16 · Shaolin Qu, Joyce Y. Chai

To tackle the vocabulary problem in conversational systems, previous work has applied unsupervised learning approaches on co-occurring speech and eye gaze during interaction to automatically acquire new words. Although t…

On the effect of curriculum learning with developmental data for grammar acquisition

2023-10-31 · Mattia Opper, J. Morrison, N. Siddharth

This work explores the degree to which grammar acquisition is driven by language `simplicity' and the source modality (speech vs. text) of data. Using BabyBERTa as a probe, we find that grammar acquisition is largely dri…

Prosodic Features from Large Corpora of Child-Directed Speech as Predictors of the Age of Acquisition of Words

2017-09-27 · Lea Frermann, Michael C. Frank

The impressive ability of children to acquire language is a widely studied phenomenon, and the factors influencing the pace and patterns of word learning remains a subject of active research. Although many models predict…

Open-source Multi-speaker Speech Corpora for Building Gujarati, Kannada, Malayalam, Marathi, Tamil and Telugu Speech Synthesis Systems

2020-05-01 · LREC 2020 5 · Fei He, Shan-Hui Cathy Chu, Oddur Kjartansson, Clara Rivera 외

We present free high quality multi-speaker speech corpora for Gujarati, Kannada, Malayalam, Marathi, Tamil and Telugu, which are six of the twenty two official languages of India spoken by 374 million native speakers. Th…

Speech Synthesistext-to-speechText to Speech

BabySLM: language-acquisition-friendly benchmark of self-supervised spoken language models

2023-06-02 · Marvin Lavechin, Yaya Sy, Hadrien Titeux, María Andrea Cruz Blandón 외

Self-supervised techniques for learning speech representations have been shown to develop linguistic competence from exposure to speech without the need for human labels. In order to fully realize the potential of these …

BenchmarkingLanguage Acquisition