MyVoice: Arabic Speech Resource Collaboration Platform
We introduce MyVoice, a crowdsourcing platform designed to collect Arabic speech to enhance dialectal speech technologies. This platform offers an opportunity to design large dialectal speech datasets; and makes them publicly available. MyVoice allows contributors to select city/country-level fine-grained dialect and record the displayed utterances. Users can switch roles between contributors and annotators. The platform incorporates a quality assurance system that filters out low-quality and spurious recordings before sending them for validation. During the validation phase, contributors can assess the quality of recordings, annotate them, and provide feedback which is then reviewed by administrators. Furthermore, the platform offers flexibility to admin roles to add new data or tasks beyond dialectal speech and word collection, which are displayed to contributors. Thus, enabling collaborative efforts in gathering diverse and large Arabic speech data.
Code (0)
등록된 구현이 없습니다.
Similar Papers 제목 키워드 기반
Navigating Dialectal Bias and Ethical Complexities in Levantine Arabic Hate Speech Detection
Social media platforms have become central to global communication, yet they also facilitate the spread of hate speech. For underrepresented dialects like Levantine Arabic, detecting hate speech presents unique cultural,…
Hate Speech DetectionOSIAN: Open Source International Arabic News Corpus - Preparation and Integration into the CLARIN-infrastructure
The World Wide Web has become a fundamental resource for building large text corpora. Broadcasting platforms such as news websites are rich sources of data regarding diverse topics and form a valuable foundation for rese…
ArticlesDescriptiveLEMMAArabicDialectHub: A Cross-Dialectal Arabic Learning Resource and Platform
We present ArabicDialectHub, a cross-dialectal Arabic learning resource comprising 552 phrases across six varieties (Moroccan Darija, Lebanese, Syrian, Emirati, Saudi, and MSA) and an interactive web platform. Phrases we…
Distractor GenerationLLM-to-Speech: A Synthetic Data Pipeline for Training Dialectal Text-to-Speech Models
Despite the advances in neural text to speech (TTS), many Arabic dialectal varieties remain marginally addressed, with most resources concentrated on Modern Spoken Arabic (MSA) and Gulf dialects, leaving Egyptian Arabic …
Synthetic Data GenerationSpeaker DiarizationSpeech SynthesisText to SpeechDevelopment of a TV Broadcasts Speech Recognition System for Qatari Arabic
A major problem with dialectal Arabic speech recognition is due to the sparsity of speech resources. In this paper, a transfer learning framework is proposed to jointly use a large amount of Modern Standard Arabic (MSA) …
Arabic Speech RecognitionLanguage ModelingLanguage Modellingspeech-recognition+2