paper-with-me

Papers

Identifying Nuanced Dialect for Arabic Tweets with Deep Learning and Reverse Translation Corpus Extension System

2020-12-01 · COLING (WANLP) 2020 12 · Rawan Tahssin, Youssef Kishk, Marwan Torki

In this paper, we present our work for the NADI Shared Task (Abdul-Mageed and Habash, 2020): Nuanced Arabic Dialect Identification for Subtask-1: country-level dialect identification. We introduce a Reverse Translation Corpus Extension Systems (RTCES) to handle data imbalance along with reported results on several experimented approaches of word and document representations and different models architectures. The top scoring model was based on AraBERT (Antoun et al., 2020), with our modified extended corpus based on reverse translation of the given Arabic tweets. The selected system achieved a macro average F1 score of 20.34% on the test set, which places us as the 7th out of 18 teams in the final ranking Leaderboard.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Tasks

Dialect IdentificationTranslation

Similar Papers 제목 키워드 기반

Faheem at NADI shared task: Identifying the dialect of Arabic tweet

2020-12-01 · COLING (WANLP) 2020 12 · Nouf AlShenaifi, Aqil Azmi

This paper describes Faheem (adj. of understand), our submission to NADI (Nuanced Arabic Dialect Identification) shared task. With so many Arabic dialects being under-studied due to the scarcity of the resources, the obj…

Dialect Identificationregression

Multi-Dialect Arabic BERT for Country-Level Dialect Identification

2020-07-10 · COLING (WANLP) 2020 12 · Bashar Talafha, Mohammad Ali, Muhy Eddin Za'ter, Haitham Seelawi 외

Arabic dialect identification is a complex problem for a number of inherent properties of the language itself. In this paper, we present the experiments conducted, and the models developed by our competing team, Mawdoo3 …

Dialect IdentificationLanguage ModelingLanguage Modelling

Dialect Identification in Nuanced Arabic Tweets Using Farasa Segmentation and AraBERT

2021-02-19 · EACL (WANLP) 2021 4 · Anshul Wadhawan

This paper presents our approach to address the EACL WANLP-2021 Shared Task 1: Nuanced Arabic Dialect Identification (NADI). The task is aimed at developing a system that identifies the geographical location(country/prov…

Dialect Identification

Finding Romanized Arabic Dialect in Code-Mixed Tweets

2014-05-01 · LREC 2014 5 · Clare Voss, Stephen Tratz, Jamal Laoudi, Douglas Briesch

Recent computational work on Arabic dialect identification has focused primarily on building and annotating corpora written in Arabic script. Arabic dialects however also appear written in Roman script, especially in soc…

Dialect IdentificationLanguage Identification

QADI: Arabic Dialect Identification in the Wild

2021-04-01 · EACL (WANLP) 2021 4 · Ahmed Abdelali, Hamdy Mubarak, Younes Samih, Sabit Hassan 외

Proper dialect identification is important for a variety of Arabic NLP applications. In this paper, we present a method for rapidly constructing a tweet dataset containing a wide range of country-level Arabic dialects —c…

Dialect Identification