paper-with-me

Papers

Dialect Identification in Nuanced Arabic Tweets Using Farasa Segmentation and AraBERT

2021-02-19 · EACL (WANLP) 2021 4 · Anshul Wadhawan

This paper presents our approach to address the EACL WANLP-2021 Shared Task 1: Nuanced Arabic Dialect Identification (NADI). The task is aimed at developing a system that identifies the geographical location(country/province) from where an Arabic tweet in the form of modern standard Arabic or dialect comes from. We solve the task in two parts. The first part involves pre-processing the provided dataset by cleaning, adding and segmenting various parts of the text. This is followed by carrying out experiments with different versions of two Transformer based models, AraBERT and AraELECTRA. Our final approach achieved macro F1-scores of 0.216, 0.235, 0.054, and 0.043 in the four subtasks, and we were ranked second in MSA identification subtasks and fourth in DA identification subtasks.

📄 PDF Abstract BibTeX arXiv:2102.09749

Code (0)

등록된 구현이 없습니다.

Tasks

Dialect Identification

Methods 이 논문이 사용한 방법론

Linear Layer A Linear Layer is a projection $\mathbf{XW + b}$.
Absolute Position Encodings Absolute Position Encodings are a type of position embeddings for [Transformer-based models] where positional encodings are…
Position-Wise Feed-Forward Layer 설명 없음
Adam 설명 없음
Multi-Head Attention 설명 없음
BPE Byte Pair Encoding, or BPE, is a subword segmentation algorithm that encodes rare and unknown words as sequences of subword units. The intuition is that various word…
Residual Connection 설명 없음
Dropout Dropout is a regularization technique for neural networks that drops a unit (along with connections) at training time with a specified probability $p$ (a common value is…

Similar Papers 제목 키워드 기반

Identifying Nuanced Dialect for Arabic Tweets with Deep Learning and Reverse Translation Corpus Extension System

2020-12-01 · COLING (WANLP) 2020 12 · Rawan Tahssin, Youssef Kishk, Marwan Torki

In this paper, we present our work for the NADI Shared Task (Abdul-Mageed and Habash, 2020): Nuanced Arabic Dialect Identification for Subtask-1: country-level dialect identification. We introduce a Reverse Translation C…

Dialect IdentificationTranslation

Multi-Dialect Arabic BERT for Country-Level Dialect Identification

2020-07-10 · COLING (WANLP) 2020 12 · Bashar Talafha, Mohammad Ali, Muhy Eddin Za'ter, Haitham Seelawi 외

Arabic dialect identification is a complex problem for a number of inherent properties of the language itself. In this paper, we present the experiments conducted, and the models developed by our competing team, Mawdoo3 …

Dialect IdentificationLanguage ModelingLanguage Modelling

Faheem at NADI shared task: Identifying the dialect of Arabic tweet

2020-12-01 · COLING (WANLP) 2020 12 · Nouf AlShenaifi, Aqil Azmi

This paper describes Faheem (adj. of understand), our submission to NADI (Nuanced Arabic Dialect Identification) shared task. With so many Arabic dialects being under-studied due to the scarcity of the resources, the obj…

Dialect Identificationregression

QADI: Arabic Dialect Identification in the Wild

2021-04-01 · EACL (WANLP) 2021 4 · Ahmed Abdelali, Hamdy Mubarak, Younes Samih, Sabit Hassan 외

Proper dialect identification is important for a variety of Arabic NLP applications. In this paper, we present a method for rapidly constructing a tweet dataset containing a wide range of country-level Arabic dialects —c…

Dialect Identification

NADI 2024: The Fifth Nuanced Arabic Dialect Identification Shared Task

2024-07-06 · Muhammad Abdul-Mageed, Amr Keleg, AbdelRahim Elmadany, Chiyu Zhang 외

We describe the findings of the fifth Nuanced Arabic Dialect Identification Shared Task (NADI 2024). NADI's objective is to help advance SoTA Arabic NLP by providing guidance, datasets, modeling opportunities, and standa…

Dialect IdentificationMachine TranslationTranslationvalid