paper-with-me

Papers

Arabic Dialect Identification in the Wild

2020-05-13 · Ahmed Abdelali, Hamdy Mubarak, Younes Samih, Sabit Hassan, Kareem Darwish

We present QADI, an automatically collected dataset of tweets belonging to a wide range of country-level Arabic dialects -covering 18 different countries in the Middle East and North Africa region. Our method for building this dataset relies on applying multiple filters to identify users who belong to different countries based on their account descriptions and to eliminate tweets that are either written in Modern Standard Arabic or contain inappropriate language. The resultant dataset contains 540k tweets from 2,525 users who are evenly distributed across 18 Arab countries. Using intrinsic evaluation, we show that the labels of a set of randomly selected tweets are 91.5% accurate. For extrinsic evaluation, we are able to build effective country-level dialect identification on tweets with a macro-averaged F1-score of 60.6% across 18 classes.

📄 PDF Abstract BibTeX arXiv:2005.06557

Code (0)

등록된 구현이 없습니다.

Tasks

Dialect Identification

Similar Papers 제목 키워드 기반

Speech Recognition Challenge in the Wild: Arabic MGB-3

2017-09-21 · Ahmed Ali, Stephan Vogel, Steve Renals

This paper describes the Arabic MGB-3 Challenge - Arabic Speech Recognition in the Wild. Unlike last year's Arabic MGB-2 Challenge, for which the recognition task was based on more than 1,200 hours broadcast TV news reco…

Arabic Speech RecognitionDialect Identificationspeech-recognitionSpeech Recognition

QADI: Arabic Dialect Identification in the Wild

2021-04-01 · EACL (WANLP) 2021 4 · Ahmed Abdelali, Hamdy Mubarak, Younes Samih, Sabit Hassan 외

Proper dialect identification is important for a variety of Arabic NLP applications. In this paper, we present a method for rapidly constructing a tweet dataset containing a wide range of country-level Arabic dialects —c…

Dialect Identification

Automatic Arabic Dialect Identification Systems for Written Texts: A Survey

2020-09-26 · Maha J. Althobaiti

Arabic dialect identification is a specific task of natural language processing, aiming to automatically predict the Arabic dialect of a given text. Arabic dialect identification is the first step in various natural lang…

Dialect IdentificationMachine TranslationSentenceSpeech Synthesis+6

Automatic Dialect Detection in Arabic Broadcast Speech

2015-09-23 · Ahmed Ali, Najim Dehak, Patrick Cardinal, Sameer Khurana 외

We investigate different approaches for dialect identification in Arabic broadcast speech, using phonetic, lexical features obtained from a speech recognition system, and acoustic features using the i-vector framework. W…

Dialect IdentificationLanguage Identificationspeech-recognitionSpeech Recognition+1

The MADAR Shared Task on Arabic Fine-Grained Dialect Identification

2019-08-01 · WS 2019 8 · Houda Bouamor, Sabit Hassan, Nizar Habash

In this paper, we present the results and findings of the MADAR Shared Task on Arabic Fine-Grained Dialect Identification. This shared task was organized as part of The Fourth Arabic Natural Language Processing Workshop,…

Dialect Identification