paper-with-me

홈 › Papers

Curras + Baladi: Towards a Levantine Corpus

2022-05-19 · LREC 2022 6 · Karim El Haff, Mustafa Jarrar, Tymaa Hammouda, Fadi Zaraket

The processing of the Arabic language is a complex field of research. This is due to many factors, including the complex and rich morphology of Arabic, its high degree of ambiguity, and the presence of several regional varieties that need to be processed while taking into account their unique characteristics. When its dialects are taken into account, this language pushes the limits of NLP to find solutions to problems posed by its inherent nature. It is a diglossic language; the standard language is used in formal settings and in education and is quite different from the vernacular languages spoken in the different regions and influenced by older languages that were historically spoken in those regions. This should encourage NLP specialists to create dialect-specific corpora such as the Palestinian morphologically annotated Curras corpus of Birzeit University. In this work, we present the Lebanese Corpus Baladi that consists of around 9.6K morphologically annotated tokens. Since Lebanese and Palestinian dialects are part of the same Levantine dialectal continuum, and thus highly mutually intelligible, our proposed corpus was constructed to be used to (1) enrich Curras and transform it into a more general Levantine corpus and (2) improve Curras by solving detected errors.

📄 PDF Abstract BibTeX arXiv:2205.09692

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Shami: A Corpus of Levantine Arabic Dialects

2018-05-01 · LREC 2018 5 · Kathrein Abu Kwaik, Motaz Saad, Stergios Chatzikyriakidis, Simon Dobnik
Language Identification

A Multi-Dialect, Multi-Genre Corpus of Informal Written Arabic

2014-05-01 · LREC 2014 5 · Ryan Cotterell, Chris Callison-Burch

This paper presents a multi-dialect, multi-genre, human annotated corpus of dialectal Arabic. We collected utterances in five Arabic dialects: Levantine, Gulf, Egyptian, Iraqi and Maghrebi. We scraped newspaper websites …

Dialect Identification

YouDACC: the Youtube Dialectal Arabic Comment Corpus

2014-05-01 · LREC 2014 5 · Ahmed Salama, Houda Bouamor, Behrang Mohit, Kemal Oflazer

This paper presents YOUDACC, an automatically annotated large-scale multi-dialectal Arabic corpus collected from user comments on Youtube videos. Our corpus covers different groups of dialects: Egyptian (EG), Gulf (GU), …

Nabra: Syrian Arabic Dialects with Morphological Annotations

2023-10-26 · Amal Nayouf, Tymaa Hammouda, Mustafa Jarrar, Fadi Zaraket 외

This paper presents Nabra, a corpora of Syrian Arabic dialects with morphological annotations. A team of Syrian natives collected more than 6K sentences containing about 60K words from several sources including social me…

Sentence

Building a Corpus of Qatari Arabic Expressions

2020-05-01 · LREC 2020 5 · Sara Al-Mulla, Wajdi Zaghouani

The current Arabic natural language processing resources are mainly build to address the Modern Standard Arabic (MSA), while we witnessed some scattered efforts to build resources for various Arabic dialects such as the …