paper-with-me

홈 › Papers

Automatic Difficulty Classification of Arabic Sentences

2021-03-07 · EACL (WANLP) 2021 4 · Nouran Khallaf, Serge Sharoff

In this paper, we present a Modern Standard Arabic (MSA) Sentence difficulty classifier, which predicts the difficulty of sentences for language learners using either the CEFR proficiency levels or the binary classification as simple or complex. We compare the use of sentence embeddings of different kinds (fastText, mBERT , XLM-R and Arabic-BERT), as well as traditional language features such as POS tags, dependency trees, readability scores and frequency lists for language learners. Our best results have been achieved using fined-tuned Arabic-BERT. The accuracy of our 3-way CEFR classification is F-1 of 0.80 and 0.75 for Arabic-Bert and XLM-R classification respectively and 0.71 Spearman correlation for regression. Our binary difficulty classifier reaches F-1 0.94 and F-1 0.98 for sentence-pair semantic similarity classifier.

📄 PDF Abstract BibTeX arXiv:2103.04386

Code (0)

등록된 구현이 없습니다.

Tasks

Binary ClassificationClassificationGeneral ClassificationPOSSemantic SimilaritySemantic Textual SimilaritySentenceSentence EmbeddingsXLM-R

Methods 이 논문이 사용한 방법론

XLM-R XLM-R
mBERT mBERT

Similar Papers 제목 키워드 기반

Curriculum Learning and Pseudo-Labeling Improve the Generalization of Multi-Label Arabic Dialect Identification Models

2026-02-12 · Ali Mekky, Mohamed El Zeftawy, Lara Hassan, Amr Keleg 외 arxiv

Being modeled as a single-label classification task for a long time, recent work has argued that Arabic Dialect Identification (ADI) should be framed as a multi-label classification task. However, ADI remains constrained…

Multi-Label Classification

Towards Arabic Sentence Simplification via Classification and Generative Approaches

2022-04-20 · Nouran Khallaf, Serge Sharoff

This paper presents an attempt to build a Modern Standard Arabic (MSA) sentence-level simplification system. We experimented with sentence simplification using two approaches: (i) a classification approach leading to lex…

ClassificationLexical SimplificationSentenceWord Embeddings

MICHAEL: Mining Character-level Patterns for Arabic Dialect Identification (MADAR Challenge)

2019-08-01 · WS 2019 8 · Dhaou Ghoul, Ga{\"e}l Lejeune

We present MICHAEL, a simple lightweight method for automatic Arabic Dialect Identification on the MADAR travel domain Dialect Identification (DID). MICHAEL uses simple character-level features in order to perform a pre-…

Dialect IdentificationGeneral Classification

Estimating the Level of Dialectness Predicts Interannotator Agreement in Multi-dialect Arabic Datasets

2024-05-18 · Amr Keleg, Walid Magdy, Sharon Goldwater

On annotating multi-dialect Arabic datasets, it is common to randomly assign the samples across a pool of native Arabic speakers. Recent analyses recommended routing dialectal samples to native speakers of their respecti…

SentenceSentence Classification

Sentiment Analysis For Modern Standard Arabic And Colloquial

2015-05-12 · Hossam S. Ibrahim, Sherif M. Abdou, Mervat Gheith

The rise of social media such as blogs and social networks has fueled interest in sentiment analysis. With the proliferation of reviews, ratings, recommendations and other forms of online expression, online opinion has t…

Arabic Sentiment AnalysisNegationSentenceSentiment Analysis+1