paper-with-me

Papers

Improved low-resource Somali speech recognition by semi-supervised acoustic and language model training

2019-07-06 · Astik Biswas, Raghav Menon, Ewald van der Westhuizen, Thomas Niesler

We present improvements in automatic speech recognition (ASR) for Somali, a currently extremely under-resourced language. This forms part of a continuing United Nations (UN) effort to employ ASR-based keyword spotting systems to support humanitarian relief programmes in rural Africa. Using just 1.57 hours of annotated speech data as a seed corpus, we increase the pool of training data by applying semi-supervised training to 17.55 hours of untranscribed speech. We make use of factorised time-delay neural networks (TDNN-F) for acoustic modelling, since these have recently been shown to be effective in resource-scarce situations. Three semi-supervised training passes were performed, where the decoded output from each pass was used for acoustic model training in the subsequent pass. The automatic transcriptions from the best performing pass were used for language model augmentation. To ensure the quality of automatic transcriptions, decoder confidence is used as a threshold. The acoustic and language models obtained from the semi-supervised approach show significant improvement in terms of WER and perplexity compared to the baseline. Incorporating the automatically generated transcriptions yields a 6.55\% improvement in language model perplexity. The use of 17.55 hour of Somali acoustic data in semi-supervised training shows an improvement of 7.74\% relative over the baseline.

📄 PDF Abstract BibTeX arXiv:1907.03064

Code (0)

등록된 구현이 없습니다.

Tasks

Acoustic ModellingAutomatic Speech RecognitionAutomatic Speech Recognition (ASR)DecoderHumanitarianKeyword SpottingLanguage ModelingLanguage Modellingspeech-recognitionSpeech Recognition

Similar Papers 제목 키워드 기반

Automatic Speech Recognition for Humanitarian Applications in Somali

2018-07-23 · Raghav Menon, Astik Biswas, Armin Saeb, John Quinn 외

We present our first efforts in building an automatic speech recognition system for Somali, an under-resourced language, using 1.57 hrs of annotated speech for acoustic model training. The system is part of an ongoing ef…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Data AugmentationHumanitarian+5

Fast Development of ASR in African Languages using Self Supervised Speech Representation Learning

2021-03-16 · Jama Hussein Mohamud, Lloyd Acquaye Thompson, Aissatou Ndoye, Laurent Besacier

This paper describes the results of an informal collaboration launched during the African Master of Machine Intelligence (AMMI) in June 2020. After a series of lectures and labs on speech data collection using mobile app…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Representation Learningspeech-recognition+2

Improving noisy student training for low-resource languages in End-to-End ASR using CycleGAN and inter-domain losses

2024-07-26 · Chia-Yu Li, Ngoc Thang Vu

Training a semi-supervised end-to-end speech recognition system using noisy student training has significantly improved performance. However, this approach requires a substantial amount of paired speech-text and unlabele…

Automatic Speech Recognitionspeech-recognitionSpeech Recognition

AfriVoices-KE: A Multilingual Speech Dataset for Kenyan Languages

2026-04-09 · Lilian Wanzare, Cynthia Amol, Ezekiel Maina, Nelson Odhiambo 외 arxiv

AfriVoices-KE is a large-scale multilingual speech dataset comprising approximately 3,000 hours of audio across five Kenyan languages: Dholuo, Kikuyu, Kalenjin, Maasai, and Somali. The dataset includes 750 hours of scrip…

Speech Recognition

Large scale weakly and semi-supervised learning for low-resource video ASR

2020-05-16 · Kritika Singh, Vimal Manohar, Alex Xiao, Sergey Edunov 외

Many semi- and weakly-supervised approaches have been investigated for overcoming the labeling cost of building high quality speech recognition systems. On the challenging task of transcribing social media videos in low-…

Decoderspeech-recognitionSpeech Recognition