MICHAEL: Mining Character-level Patterns for Arabic Dialect Identification (MADAR Challenge)
We present MICHAEL, a simple lightweight method for automatic Arabic Dialect Identification on the MADAR travel domain Dialect Identification (DID). MICHAEL uses simple character-level features in order to perform a pre-processing free classification. More precisely, Character N-grams extracted from the original sentences are used to train a Multinomial Naive Bayes classifier. This system achieved an official score (accuracy) of 53.25{\%} with 1{\textless}=N{\textless}=3 but showed a much better result with character 4-grams (62.17{\%} accuracy).
Code (0)
등록된 구현이 없습니다.
Tasks
Dialect IdentificationGeneral ClassificationSimilar Papers 제목 키워드 기반
Arabizi Detection and Conversion to Arabic
Arabizi is Arabic text that is written using Latin characters. Arabizi is used to present both Modern Standard Arabic (MSA) or Arabic dialects. It is commonly used in informal settings such as social networking sites and…
Language ModelingLanguage ModellingTransliterationGeometric Patterns of Meaning: A PHATE Manifold Analysis of Multi-lingual Embeddings
We introduce a multi-level analysis framework for examining semantic geometry in multilingual embeddings, implemented through Semanscope (a visualization tool that applies PHATE manifold learning across four linguistic l…
EMA at SemEval-2018 Task 1: Emotion Mining for Arabic
While significant progress has been achieved for Opinion Mining in Arabic (OMA), very limited efforts have been put towards the task of Emotion mining in Arabic. In fact, businesses are interested in learning a fine-grai…
Emotion ClassificationEmotion RecognitionOpinion MiningRecommendation Systems+2A Characterization Study of Arabic Twitter Data with a Benchmarking for State-of-the-Art Opinion Mining Models
Opinion mining in Arabic is a challenging task given the rich morphology of the language. The task becomes more challenging when it is applied to Twitter data, which contains additional sources of noise, such as the use …
BenchmarkingFeature EngineeringOpinion MiningSentiment AnalysisI3rab: A New Arabic Dependency Treebank Based on Arabic Grammatical Theory
Treebanks are valuable linguistic resources that include the syntactic structure of a language sentence in addition to POS-tags and morphological features. They are mainly utilized in modeling statistical parsers. Althou…
POSSentence