paper-with-me

홈 › Papers

QADI: Arabic Dialect Identification in the Wild

2021-04-01 · EACL (WANLP) 2021 4 · Ahmed Abdelali, Hamdy Mubarak, Younes Samih, Sabit Hassan, Kareem Darwish

Proper dialect identification is important for a variety of Arabic NLP applications. In this paper, we present a method for rapidly constructing a tweet dataset containing a wide range of country-level Arabic dialects —covering 18 different countries in the Middle East and North Africa region. Our method relies on applying multiple filters to identify users who belong to different countries based on their account descriptions and to eliminate tweets that either write mainly in Modern Standard Arabic or mostly use vulgar language. The resultant dataset contains 540k tweets from 2,525 users who are evenly distributed across 18 Arab countries. Using intrinsic evaluation, we show that the labels of a set of randomly selected tweets are 91.5% accurate. For extrinsic evaluation, we are able to build effective country level dialect identification on tweets with a macro-averaged F1-score of 60.6% across 18 classes.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Tasks

Dialect Identification

Similar Papers 제목 키워드 기반

Arabic Dialect Identification in the Wild

2020-05-13 · Ahmed Abdelali, Hamdy Mubarak, Younes Samih, Sabit Hassan 외

We present QADI, an automatically collected dataset of tweets belonging to a wide range of country-level Arabic dialects -covering 18 different countries in the Middle East and North Africa region. Our method for buildin…

Dialect Identification

Speech Recognition Challenge in the Wild: Arabic MGB-3

2017-09-21 · Ahmed Ali, Stephan Vogel, Steve Renals

This paper describes the Arabic MGB-3 Challenge - Arabic Speech Recognition in the Wild. Unlike last year's Arabic MGB-2 Challenge, for which the recognition task was based on more than 1,200 hours broadcast TV news reco…

Arabic Speech RecognitionDialect Identificationspeech-recognitionSpeech Recognition

Computational Linguistics Meets Libyan Dialect: A Study on Dialect Identification

2025-12-03 · Mansour Essgaer, Khamis Massud, Rabia Al Mamlook, Najah Ghmaid arxiv

This study investigates logistic regression, linear support vector machine, multinomial Naive Bayes, and Bernoulli Naive Bayes for classifying Libyan dialect utterances gathered from Twitter. The dataset used is the QADI…

Automatic Arabic Dialect Identification Systems for Written Texts: A Survey

2020-09-26 · Maha J. Althobaiti

Arabic dialect identification is a specific task of natural language processing, aiming to automatically predict the Arabic dialect of a given text. Arabic dialect identification is the first step in various natural lang…

Dialect IdentificationMachine TranslationSentenceSpeech Synthesis+6

Automatic Dialect Detection in Arabic Broadcast Speech

2015-09-23 · Ahmed Ali, Najim Dehak, Patrick Cardinal, Sameer Khurana 외

We investigate different approaches for dialect identification in Arabic broadcast speech, using phonetic, lexical features obtained from a speech recognition system, and acoustic features using the i-vector framework. W…

Dialect IdentificationLanguage Identificationspeech-recognitionSpeech Recognition+1