Dialectal Coverage And Generalization in Arabic Speech Recognition
Developing robust automatic speech recognition (ASR) systems for Arabic requires effective strategies to manage its diversity. Existing ASR systems mainly cover the modern standard Arabic (MSA) variety and few high-resource dialects, but fall short in coverage and generalization across the multitude of spoken variants. Code-switching with English and French is also common in different regions of the Arab world, which challenges the performance of monolingual Arabic models. In this work, we introduce a suite of ASR models optimized to effectively recognize multiple variants of spoken Arabic, including MSA, various dialects, and code-switching. We provide open-source pre-trained models that cover data from 17 Arabic-speaking countries, and fine-tuned MSA and dialectal ASR models that include at least 11 variants, as well as multi-lingual ASR models covering embedded languages in code-switched utterances. We evaluate ASR performance across these spoken varieties and demonstrate both coverage and performance gains compared to prior models.
Code (1)
Tasks
Arabic Speech RecognitionAutomatic Speech RecognitionAutomatic Speech Recognition (ASR)Diversityspeech-recognitionSpeech RecognitionSimilar Papers 제목 키워드 기반
Development of a TV Broadcasts Speech Recognition System for Qatari Arabic
A major problem with dialectal Arabic speech recognition is due to the sparsity of speech resources. In this paper, a transfer learning framework is proposed to jointly use a large amount of Modern Standard Arabic (MSA) …
Arabic Speech RecognitionLanguage ModelingLanguage Modellingspeech-recognition+2Overcoming Data Scarcity in Multi-Dialectal Arabic ASR via Whisper Fine-Tuning
Although commercial Arabic automatic speech recognition (ASR) systems support Modern Standard Arabic (MSA), they struggle with dialectal speech. We investigate the effect of fine-tuning OpenAI's Whisper on five major Ara…
Automatic Speech RecognitionAutomatic Speech Recognition (ASR)speech-recognitionSpeech RecognitionAra-Best-RQ: Multi Dialectal Arabic SSL
We present Ara-BEST-RQ, a family of self-supervised learning (SSL) models specifically designed for multi-dialectal Arabic speech processing. Leveraging 5,640 hours of crawled Creative Commons speech and combining it wit…
Self-Supervised LearningSpeech RecognitionArTST: Arabic Text and Speech Transformer
We present ArTST, a pre-trained Arabic text and speech transformer for supporting open-source speech technologies for the Arabic language. The model architecture follows the unified-modal framework, SpeechT5, that was re…
Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Dialect Identificationspeech-recognition+5Towards One Model to Rule All: Multilingual Strategy for Dialectal Code-Switching Arabic ASR
With the advent of globalization, there is an increasing demand for multilingual automatic speech recognition (ASR), handling language and dialectal variation of spoken content. Recent studies show its efficacy over mono…
AllAutomatic Speech RecognitionAutomatic Speech Recognition (ASR)speech-recognition+1