paper-with-me

Papers

Sagalee: an Open Source Automatic Speech Recognition Dataset for Oromo Language

2025-02-01 · Turi Abu, Ying Shi, Thomas Fang Zheng, Dong Wang

We present a novel Automatic Speech Recognition (ASR) dataset for the Oromo language, a widely spoken language in Ethiopia and neighboring regions. The dataset was collected through a crowd-sourcing initiative, encompassing a diverse range of speakers and phonetic variations. It consists of 100 hours of real-world audio recordings paired with transcriptions, covering read speech in both clean and noisy environments. This dataset addresses the critical need for ASR resources for the Oromo language which is underrepresented. To show its applicability for the ASR task, we conducted experiments using the Conformer model, achieving a Word Error Rate (WER) of 15.32% with hybrid CTC and AED loss and WER of 18.74% with pure CTC loss. Additionally, fine-tuning the Whisper model resulted in a significantly improved WER of 10.82%. These results establish baselines for Oromo ASR, highlighting both the challenges and the potential for improving ASR performance in Oromo. The dataset is publicly available at https://github.com/turinaf/sagalee and we encourage its use for further research and development in Oromo speech processing.

📄 PDF Abstract BibTeX arXiv:2502.00421

Code (1)

turinaf/sagalee 공식 구현 pytorch

Tasks

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)speech-recognitionSpeech Recognition

Similar Papers 제목 키워드 기반

KoSpeech: Open-Source Toolkit for End-to-End Korean Speech Recognition

2020-09-07 · Soohwan Kim, Seyoung Bae, Cheolhwang Won

We present KoSpeech, an open-source software, which is modular and extensible end-to-end Korean automatic speech recognition (ASR) toolkit based on the deep learning library PyTorch. Several automatic speech recognition …

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)speech-recognitionSpeech Recognition

Bengali Common Voice Speech Dataset for Automatic Speech Recognition

2022-06-28 · Samiul Alam, Asif Sushmit, Zaowad Abdullah, Shahrin Nakkhatra 외

Bengali is one of the most spoken languages in the world with over 300 million speakers globally. Despite its popularity, research into the development of Bengali speech recognition systems is hindered due to the lack of…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)DiversitySentence+2

MooER: LLM-based Speech Recognition and Translation Models from Moore Threads

2024-08-09 · Junhao Xu, Zhenlin Liang, Yi Liu, Yichao Hu 외

In this paper, we present MooER, a LLM-based large-scale automatic speech recognition (ASR) / automatic speech translation (AST) model of Moore Threads. A 5000h pseudo labeled dataset containing open source and self coll…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)automatic-speech-translationspeech-recognition+1

MOSEL: 950,000 Hours of Speech Data for Open-Source Speech Foundation Model Training on EU Languages

2024-10-01 · Marco Gaido, Sara Papi, Luisa Bentivogli, Alessio Brutti 외

The rise of foundation models (FMs), coupled with regulatory efforts addressing their risks and impacts, has sparked significant interest in open-source models. However, existing speech FMs (SFMs) fall short of full comp…

Automatic Speech Recognitionspeech-recognitionSpeech Recognition

ESPnet: End-to-End Speech Processing Toolkit

2018-03-30 · Shinji Watanabe, Takaaki Hori, Shigeki Karita, Tomoki Hayashi 외

This paper introduces a new open source platform for end-to-end speech processing named ESPnet. ESPnet mainly focuses on end-to-end automatic speech recognition (ASR), and adopts widely-used dynamic neural network toolki…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)speech-recognitionSpeech Recognition