paper-with-me

Papers

Finding Romanized Arabic Dialect in Code-Mixed Tweets

2014-05-01 · LREC 2014 5 · Clare Voss, Stephen Tratz, Jamal Laoudi, Douglas Briesch

Recent computational work on Arabic dialect identification has focused primarily on building and annotating corpora written in Arabic script. Arabic dialects however also appear written in Roman script, especially in social media. This paper describes our recent work developing tweet corpora and a token-level classifier that identifies a Romanized Arabic dialect and distinguishes it from French and English in tweets. We focus on Moroccan Darija, one of several spoken vernaculars in the family of Maghrebi Arabic dialects. Even given noisy, code-mixed tweets,the classifier achieved token-level recall of 93.2{\%} on Romanized Arabic dialect, 83.2{\%} on English, and 90.1{\%} on French. The classifier, now integrated into our tweet conversation annotation tool (Tratz et al. 2013), has semi-automated the construction of a Romanized Arabic-dialect lexicon. Two datasets, a full list of Moroccan Darija surface token forms and a table of lexical entries derived from this list with spelling variants, as extracted from our tweet corpus collection, will be made available in the LRE MAP.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Tasks

Dialect IdentificationLanguage Identification

Similar Papers 제목 키워드 기반

Automatic Transliteration of Romanized Dialectal Arabic

2014-06-01 · WS 2014 6 · Mohamed Al-Badrashiny, Esk, Ramy er, Nizar Habash 외
Language ModellingSpelling CorrectionTransliteration

Script Sensitivity: Benchmarking Language Models on Unicode, Romanized and Mixed-Script Sinhala

2026-01-21 · Minuri Rajapakse, Ruvan Weerasinghe arxiv

The performance of Language Models (LMs) on low-resource, morphologically rich languages like Sinhala remains largely unexplored, particularly regarding script variation in digital communication. Sinhala exhibits script …

Beyond MCQ: An Open-Ended Arabic Cultural QA Benchmark with Dialect Variants

2025-10-28 · Hunzalah Hassan Bhatti, Firoj Alam arxiv

Large Language Models (LLMs) are increasingly used to answer everyday questions, yet their performance on culturally grounded and dialectal content remains uneven across languages. We propose a comprehensive method that …

Building a User-Generated Content North-African Arabizi Treebank: Tackling Hell

2020-07-01 · ACL 2020 6 · Djam{\'e} Seddah, Farah Essaidi, Amal Fethi, Matthieu Futeral 외

We introduce the first treebank for a romanized user-generated content variety of Algerian, a North-African Arabic dialect known for its frequent usage of code-switching. Made of 1500 sentences, fully annotated in morpho…

Dependency ParsingPOSPOS TaggingSentence+1

Towards One Model to Rule All: Multilingual Strategy for Dialectal Code-Switching Arabic ASR

2021-05-31 · Shammur Absar Chowdhury, Amir Hussein, Ahmed Abdelali, Ahmed Ali

With the advent of globalization, there is an increasing demand for multilingual automatic speech recognition (ASR), handling language and dialectal variation of spoken content. Recent studies show its efficacy over mono…

AllAutomatic Speech RecognitionAutomatic Speech Recognition (ASR)speech-recognition+1