paper-with-me

Papers

MANorm: A Normalization Dictionary for Moroccan Arabic Dialect Written in Latin Script

2022-06-18 · COLING (WANLP) 2020 12 · Randa Zarnoufi, Walid Bachri, Hamid Jaafar, Mounia Abik

Social media user-generated text is actually the main resource for many NLP tasks. This text however, does not follow the standard rules of writing. Moreover, the use of dialect such as Moroccan Arabic in written communications increases further NLP tasks complexity. A dialect is a verbal language that does not have a standard orthography, which leads users to improvise spelling while writing. Thus, for the same word we can find multiple forms of transliterations. Subsequently, it is mandatory to normalize these different transliterations to one canonical word form. To reach this goal, we have exploited the powerfulness of word embedding models generated with a corpus of YouTube comments. Besides, using a Moroccan Arabic dialect dictionary that provides the canonical forms, we have built a normalization dictionary that we refer to as MANorm. We have conducted several experiments to demonstrate the efficiency of MANorm, which have shown its usefulness in dialect normalization.

📄 PDF Abstract BibTeX arXiv:2206.09167

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Sentiment Analysis Dataset in Moroccan Dialect: Bridging the Gap Between Arabic and Latin Scripted dialect

2023-03-28 · Mouad Jbel, Mourad Jabrane, Imad Hafidi, Abdulmutallib Metrane

Sentiment analysis, the automated process of determining emotions or opinions expressed in text, has seen extensive exploration in the field of natural language processing. However, one aspect that has remained underrepr…

Sentiment AnalysisSentiment Classification

Morphologically Annotated Corpora and Morphological Analyzers for Moroccan and Sanaani Yemeni Arabic

2016-05-01 · LREC 2016 5 · Faisal Al-Shargi, Aidan Kaplan, Esk, Ramy er 외

We present new language resources for Moroccan and Sanaani Yemeni Arabic. The resources include corpora for each dialect which have been morphologically annotated, and morphological analyzers for each dialect which are d…

Morphologically Annotated Corpora for Seven Arabic Dialects: Taizi, Sanaani, Najdi, Jordanian, Syrian, Iraqi and Moroccan

2019-08-01 · WS 2019 8 · Faisal Alshargi, Shahd Dibas, Sakhar Alkhereyf, Reem Faraj 외

We present a collection of morphologically annotated corpora for seven Arabic dialects: Taizi Yemeni, Sanaani Yemeni, Najdi, Jordanian, Syrian, Iraqi and Moroccan Arabic. The corpora collectively cover over 200,000 words…

Morphological Analysis

Diacritization of Maghrebi Arabic Sub-Dialects

2018-10-15 · Ahmed Abdelali, Mohammed Attia, Younes Samih, Kareem Darwish 외

Diacritization process attempt to restore the short vowels in Arabic written text; which typically are omitted. This process is essential for applications such as Text-to-Speech (TTS). While diacritization of Modern Stan…

text-to-speechText to Speech

Finding Romanized Arabic Dialect in Code-Mixed Tweets

2014-05-01 · LREC 2014 5 · Clare Voss, Stephen Tratz, Jamal Laoudi, Douglas Briesch

Recent computational work on Arabic dialect identification has focused primarily on building and annotating corpora written in Arabic script. Arabic dialects however also appear written in Roman script, especially in soc…

Dialect IdentificationLanguage Identification