Automatic Difficulty Classification of Arabic Sentences
In this paper, we present a Modern Standard Arabic (MSA) Sentence difficulty classifier, which predicts the difficulty of sentences for language learners using either the CEFR proficiency levels or the binary classification as simple or complex. We compare the use of sentence embeddings of different kinds (fastText, mBERT , XLM-R and Arabic-BERT), as well as traditional language features such as POS tags, dependency trees, readability scores and frequency lists for language learners. Our best results have been achieved using fined-tuned Arabic-BERT. The accuracy of our 3-way CEFR classification is F-1 of 0.80 and 0.75 for Arabic-Bert and XLM-R classification respectively and 0.71 Spearman correlation for regression. Our binary difficulty classifier reaches F-1 0.94 and F-1 0.98 for sentence-pair semantic similarity classifier.
Code (0)
등록된 구현이 없습니다.
Tasks
Binary ClassificationClassificationGeneral ClassificationPOSSemantic SimilaritySemantic Textual SimilaritySentenceSentence EmbeddingsXLM-RMethods 이 논문이 사용한 방법론
Similar Papers 제목 키워드 기반
Curriculum Learning and Pseudo-Labeling Improve the Generalization of Multi-Label Arabic Dialect Identification Models
Being modeled as a single-label classification task for a long time, recent work has argued that Arabic Dialect Identification (ADI) should be framed as a multi-label classification task. However, ADI remains constrained…
Multi-Label ClassificationTowards Arabic Sentence Simplification via Classification and Generative Approaches
This paper presents an attempt to build a Modern Standard Arabic (MSA) sentence-level simplification system. We experimented with sentence simplification using two approaches: (i) a classification approach leading to lex…
ClassificationLexical SimplificationSentenceWord EmbeddingsMICHAEL: Mining Character-level Patterns for Arabic Dialect Identification (MADAR Challenge)
We present MICHAEL, a simple lightweight method for automatic Arabic Dialect Identification on the MADAR travel domain Dialect Identification (DID). MICHAEL uses simple character-level features in order to perform a pre-…
Dialect IdentificationGeneral ClassificationEstimating the Level of Dialectness Predicts Interannotator Agreement in Multi-dialect Arabic Datasets
On annotating multi-dialect Arabic datasets, it is common to randomly assign the samples across a pool of native Arabic speakers. Recent analyses recommended routing dialectal samples to native speakers of their respecti…
SentenceSentence ClassificationSentiment Analysis For Modern Standard Arabic And Colloquial
The rise of social media such as blogs and social networks has fueled interest in sentiment analysis. With the proliferation of reviews, ratings, recommendations and other forms of online expression, online opinion has t…
Arabic Sentiment AnalysisNegationSentenceSentiment Analysis+1