Text and Speech-based Tunisian Arabic Sub-Dialects Identification
Dialect IDentification (DID) is a challenging task, and it becomes more complicated when it is about the identification of dialects that belong to the same country. Indeed, dialects of the same country are closely related and exhibit a significant overlapping at the phonetic and lexical levels. In this paper, we present our first results on a dialect classification task covering four sub-dialects spoken in Tunisia. We use the term {'}sub-dialect{'} to refer to the dialects belonging to the same country. We conducted our experiments aiming to discriminate between Tunisian sub-dialects belonging to four different cities: namely Tunis, Sfax, Sousse and Tataouine. A spoken corpus of 1673 utterances is collected, transcribed and freely distributed. We used this corpus to build several speech- and text-based DID systems. Our results confirm that, at this level of granularity, dialects are much better distinguishable using the speech modality. Indeed, we were able to reach an F-1 score of 93.75{\%} using our best speech-based identification system while the F-1 score is limited to 54.16{\%} using text-based DID on the same test set.
Code (0)
등록된 구현이 없습니다.
Tasks
Dialect IdentificationSimilar Papers 제목 키워드 기반
Sentiment Analysis of Tunisian Dialects: Linguistic Ressources and Experiments
Dialectal Arabic (DA) is significantly different from the Arabic language taught in schools and used in written communication and formal speech (broadcast news, religion, politics, etc.). There are many existing research…
Sentiment AnalysisDiacritization of Maghrebi Arabic Sub-Dialects
Diacritization process attempt to restore the short vowels in Arabic written text; which typically are omitted. This process is essential for applications such as Text-to-Speech (TTS). While diacritization of Modern Stan…
text-to-speechText to SpeechTEDxTN: A Three-way Speech Translation Corpus for Code-Switched Tunisian Arabic - English
In this paper, we introduce TEDxTN, the first publicly available Tunisian Arabic to English speech translation dataset. This work is in line with the ongoing effort to mitigate the data scarcity obstacle for a number of …
Speech RecognitionTArC: Incrementally and Semi-Automatically Collecting a Tunisian Arabish Corpus
This article describes the constitution process of the first morpho-syntactically annotated Tunisian Arabish Corpus (TArC). Arabish, also known as Arabizi, is a spontaneous coding of Arabic dialects in Latin characters a…
Leveraging Data Collection and Unsupervised Learning for Code-switched Tunisian Arabic Automatic Speech Recognition
Crafting an effective Automatic Speech Recognition (ASR) solution for dialects demands innovative approaches that not only address the data scarcity issue but also navigate the intricacies of linguistic diversity. In thi…
Automatic Speech RecognitionAutomatic Speech Recognition (ASR)DiversityNavigate+2