paper-with-me

Papers

Text and Speech-based Tunisian Arabic Sub-Dialects Identification

2020-05-01 · LREC 2020 5 · Najla Ben Abdallah, Sam{\'e}h Kchaou, Fethi Bougares

Dialect IDentification (DID) is a challenging task, and it becomes more complicated when it is about the identification of dialects that belong to the same country. Indeed, dialects of the same country are closely related and exhibit a significant overlapping at the phonetic and lexical levels. In this paper, we present our first results on a dialect classification task covering four sub-dialects spoken in Tunisia. We use the term {'}sub-dialect{'} to refer to the dialects belonging to the same country. We conducted our experiments aiming to discriminate between Tunisian sub-dialects belonging to four different cities: namely Tunis, Sfax, Sousse and Tataouine. A spoken corpus of 1673 utterances is collected, transcribed and freely distributed. We used this corpus to build several speech- and text-based DID systems. Our results confirm that, at this level of granularity, dialects are much better distinguishable using the speech modality. Indeed, we were able to reach an F-1 score of 93.75{\%} using our best speech-based identification system while the F-1 score is limited to 54.16{\%} using text-based DID on the same test set.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Tasks

Dialect Identification

Similar Papers 제목 키워드 기반

Sentiment Analysis of Tunisian Dialects: Linguistic Ressources and Experiments

2017-04-01 · WS 2017 4 · Salima Medhaffar, Fethi Bougares, Yannick Est{\`e}ve, Lamia Hadrich-Belguith

Dialectal Arabic (DA) is significantly different from the Arabic language taught in schools and used in written communication and formal speech (broadcast news, religion, politics, etc.). There are many existing research…

Sentiment Analysis

Diacritization of Maghrebi Arabic Sub-Dialects

2018-10-15 · Ahmed Abdelali, Mohammed Attia, Younes Samih, Kareem Darwish 외

Diacritization process attempt to restore the short vowels in Arabic written text; which typically are omitted. This process is essential for applications such as Text-to-Speech (TTS). While diacritization of Modern Stan…

text-to-speechText to Speech

TEDxTN: A Three-way Speech Translation Corpus for Code-Switched Tunisian Arabic - English

2025-11-13 · Fethi Bougares, Salima Mdhaffar, Haroun Elleuch, Yannick Estève arxiv

In this paper, we introduce TEDxTN, the first publicly available Tunisian Arabic to English speech translation dataset. This work is in line with the ongoing effort to mitigate the data scarcity obstacle for a number of …

Speech Recognition

TArC: Incrementally and Semi-Automatically Collecting a Tunisian Arabish Corpus

2020-03-20 · LREC 2020 5 · Elisa Gugliotta, Marco Dinarelli

This article describes the constitution process of the first morpho-syntactically annotated Tunisian Arabish Corpus (TArC). Arabish, also known as Arabizi, is a spontaneous coding of Arabic dialects in Latin characters a…

Leveraging Data Collection and Unsupervised Learning for Code-switched Tunisian Arabic Automatic Speech Recognition

2023-09-20 · Ahmed Amine Ben Abdallah, Ata Kabboudi, Amir Kanoun, Salah Zaiem

Crafting an effective Automatic Speech Recognition (ASR) solution for dialects demands innovative approaches that not only address the data scarcity issue but also navigate the intricacies of linguistic diversity. In thi…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)DiversityNavigate+2