paper-with-me

홈 › Papers

Ramsa: A Large Sociolinguistically Rich Emirati Arabic Speech Corpus for ASR and TTS

2026-03-09 · Rania Al-Sabbagh arxiv

Ramsa is a developing 41-hour speech corpus of Emirati Arabic designed to support sociolinguistic research and low-resource language technologies. It contains recordings from structured interviews with native speakers and episodes from national television shows. The corpus features 157 speakers (59 female, 98 male), spans subdialects such as Urban, Bedouin, and Mountain/Shihhi, and covers topics such as cultural heritage, agriculture and sustainability, daily life, professional trajectories, and architecture. It consists of 91 monologic and 79 dialogic recordings, varying in length and recording conditions. A 10\% subset was used to evaluate commercial and open-source models for automatic speech recognition (ASR) and text-to-speech (TTS) in a zero-shot setting to establish initial baselines. Whisper-large-v3-turbo achieved the best ASR performance, with average word and character error rates of 0.268 and 0.144, respectively. MMS-TTS-Ara reported the best mean word and character rates of 0.285 and 0.081, respectively, for TTS. These baselines are competitive but leave substantial room for improvement. The paper highlights the challenges encountered and provides directions for future work.

📄 PDF Abstract BibTeX arXiv:2603.08125

Code (0)

등록된 구현이 없습니다.

Tasks

Speech Recognition

Similar Papers 제목 키워드 기반

Building the Emirati Arabic FrameNet

2020-05-01 · LREC 2020 5 · Andrew Gargett, Tommi Leung

The Emirati Arabic FrameNet (EAFN) project aims to initiate a FrameNet for Emirati Arabic, utilizing the Emirati Arabic Corpus. The goal is to create a resource comparable to the initial stages of the Berkeley FrameNet. …

Mixat: A Data Set of Bilingual Emirati-English Speech

2024-05-04 · Maryam Al Ali, Hanan Aldarmaki

This paper introduces Mixat: a dataset of Emirati speech code-mixed with English. Mixat was developed to address the shortcomings of current speech recognition resources when applied to Emirati speech, and in particular,…

speech-recognitionSpeech Recognition

A Morphologically Annotated Corpus of Emirati Arabic

2018-05-01 · LREC 2018 5 · Salam Khalifa, Nizar Habash, Fadhl Eryani, Ossama Obeid 외
LemmatizationMachine TranslationMorphological AnalysisPart-Of-Speech Tagging

ArabicDialectHub: A Cross-Dialectal Arabic Learning Resource and Platform

2026-01-30 · Salem Lahlou arxiv

We present ArabicDialectHub, a cross-dialectal Arabic learning resource comprising 552 phrases across six varieties (Moroccan Darija, Lebanese, Syrian, Emirati, Saudi, and MSA) and an interactive web platform. Phrases we…

Distractor Generation

Casablanca: Data and Models for Multidialectal Arabic Speech Recognition

2024-10-06 · Bashar Talafha, Karima Kadaoui, Samar Mohamed Magdy, Mariem Habiboullah 외

In spite of the recent progress in speech processing, the majority of world languages and dialects remain uncovered. This situation only furthers an already wide technological divide, thereby hindering technological and …

Arabic Speech Recognitionspeech-recognitionSpeech Recognition