paper-with-me

홈 › Papers

Building a Luganda Text-to-Speech Model From Crowdsourced Data

2024-05-16 · Sulaiman Kagumire, Andrew Katumba, Joyce Nakatumba-Nabende, John Quinn

Text-to-speech (TTS) development for African languages such as Luganda is still limited, primarily due to the scarcity of high-quality, single-speaker recordings essential for training TTS models. Prior work has focused on utilizing the Luganda Common Voice recordings of multiple speakers aged between 20-49. Although the generated speech is intelligible, it is still of lower quality than the model trained on studio-grade recordings. This is due to the insufficient data preprocessing methods applied to improve the quality of the Common Voice recordings. Furthermore, speech convergence is more difficult to achieve due to varying intonations, as well as background noise. In this paper, we show that the quality of Luganda TTS from Common Voice can improve by training on multiple speakers of close intonation in addition to further preprocessing of the training data. Specifically, we selected six female speakers with close intonation determined by subjectively listening and comparing their voice recordings. In addition to trimming out silent portions from the beginning and end of the recordings, we applied a pre-trained speech enhancement model to reduce background noise and enhance audio quality. We also utilized a pre-trained, non-intrusive, self-supervised Mean Opinion Score (MOS) estimation model to filter recordings with an estimated MOS over 3.5, indicating high perceived quality. Subjective MOS evaluations from nine native Luganda speakers demonstrate that our TTS model achieves a significantly better MOS of 3.55 compared to the reported 2.5 MOS of the existing model. Moreover, for a fair comparison, our model trained on six speakers outperforms models trained on a single-speaker (3.13 MOS) or two speakers (3.22 MOS). This showcases the effectiveness of compensating for the lack of data from one speaker with data from multiple speakers of close intonation to improve TTS quality.

📄 PDF Abstract BibTeX arXiv:2405.10211

Code (0)

등록된 구현이 없습니다.

Tasks

Speech Enhancementtext-to-speechText to Speech

Similar Papers 제목 키워드 기반

The Makerere Radio Speech Corpus: A Luganda Radio Corpus for Automatic Speech Recognition

2022-06-20 · LREC 2022 6 · Jonathan Mukiibi, Andrew Katumba, Joyce Nakatumba-Nabende, Ali Hussein 외

Building a usable radio monitoring automatic speech recognition (ASR) system is a challenging task for under-resourced languages and yet this is paramount in societies where radio is the main medium of public communicati…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)speech-recognitionSpeech Recognition

Luganda Text-to-Speech Machine

2020-05-11 · Irene Nandutu, Ernest Mwebaze

In Uganda, Luganda is the most spoken native language. It is used for communication in informal as well as formal business transactions. The development of technology startups globally related to TTS has mainly been with…

text-to-speechText to Speech

Gender bias Evaluation in Luganda-English Machine Translation

2022-09-01 · AMTA 2022 9 · Eric Peter Wairagala

We have seen significant growth in the area of building Natural Language Processing (NLP) tools for African languages. However, the evaluation of gender bias in the machine translation systems for African languages is no…

Embeddings EvaluationFairnessMachine TranslationTransfer Learning+2

Luganda Speech Intent Recognition for IoT Applications

2024-05-16 · Andrew Katumba, Sudi Murindanyi, John Trevor Kasule, Elvis Mugume

The advent of Internet of Things (IoT) technology has generated massive interest in voice-controlled smart homes. While many voice-controlled smart home systems are designed to understand and support widely spoken langua…

intent-classificationIntent ClassificationIntent RecognitionSpeech Intent Classification

Building a Parallel Corpus and Training Translation Models Between Luganda and English

2023-01-07 · Richard Kimera, Daniela N. Rim, Heeyoul Choi

Neural machine translation (NMT) has achieved great successes with large datasets, so NMT is more premised on high-resource languages. This continuously underpins the low resource languages such as Luganda due to the lac…

Machine TranslationNMTTranslation