Towards Responsible Natural Language Annotation for the Varieties of Arabic
When building NLP models, there is a tendency to aim for broader coverage, often overlooking cultural and (socio)linguistic nuance. In this position paper, we make the case for care and attention to such nuances, particularly in dataset annotation, as well as the inclusion of cultural and linguistic expertise in the process. We present a playbook for responsible dataset creation for polyglossic, multidialectal languages. This work is informed by a study on Arabic annotation of social media content.
Code (0)
등록된 구현이 없습니다.
Tasks
PositionSimilar Papers 제목 키워드 기반
Towards Responsible Natural Language Annotation for the Varieties of Arabic
When building NLP models, there is a tendency to aim for broader coverage, often overlooking cultural and (socio)linguistic nuance. In this position paper, we make the case for care and attention to such nuances, particu…
PositionArabic natural language processing: An overview
Arabic is recognised as the 4th most used language of the Internet. Arabic has three main varieties: (1) classical Arabic (CA), (2) Modern Standard Arabic (MSA), (3) Arabic Dialect (AD). MSA and AD could be written eithe…
SurveyA Large Scale Corpus of Gulf Arabic
Most Arabic natural language processing tools and resources are developed to serve Modern Standard Arabic (MSA), which is the official written language in the Arab World. Some Dialectal Arabic varieties, notably Egyptian…
A Spelling Correction Corpus for Multiple Arabic Dialects
Arabic dialects are the non-standard varieties of Arabic commonly spoken {--} and increasingly written on social media {--} across the Arab world. Arabic dialects do not have standard orthographies, a challenge for natur…
Spelling CorrectionText NormalizationAutomatic Detection of Arabicized Berber and Arabic Varieties
Automatic Language Identification (ALI) is the detection of the natural language of an input text by a machine. It is the first necessary step to do any language-dependent natural language processing task. Various method…
Language Identification