paper-with-me

홈 › Papers

Efficient Measuring of Readability to Improve Documents Accessibility for Arabic Language Learners

2021-09-09 · Sadik Bessou, Ghozlane Chenni

This paper presents an approach based on supervised machine learning methods to build a classifier that can identify text complexity in order to present Arabic language learners with texts suitable to their levels. The approach is based on machine learning classification methods to discriminate between the different levels of difficulty in reading and understanding a text. Several models were trained on a large corpus mined from online Arabic websites and manually annotated. The model uses both Count and TF-IDF representations and applies five machine learning algorithms; Multinomial Naive Bayes, Bernoulli Naive Bayes, Logistic Regression, Support Vector Machine and Random Forest, using unigrams and bigrams features. With the goal of extracting the text complexity, the problem is usually addressed by formulating the level identification as a classification task. Experimental results showed that n-gram features could be indicative of the reading level of a text and could substantially improve performance, and showed that SVM and Multinomial Naive Bayes are the most accurate in predicting the complexity level. Best results were achieved using TF-IDF Vectors trained by a combination of word-based unigrams and bigrams with an overall accuracy of 87.14% over four classes of complexity.

📄 PDF Abstract BibTeX arXiv:2109.08648

Code (0)

등록된 구현이 없습니다.

Tasks

BIG-bench Machine Learning

Methods 이 논문이 사용한 방법론

Logistic Regression Logistic Regression, despite its name, is a linear model for classification rather than regression. Logistic regression is also known in the literature as logit regression,…
SVM A Support Vector Machine, or SVM, is a non-parametric supervised learning model. For non-linear classification and regression, they utilise the kernel trick to map inputs…

Similar Papers 제목 키워드 기반

OSMAN ― A Novel Arabic Readability Metric

2016-05-01 · LREC 2016 5 · Mahmoud El-Haj, Paul Rayson

We present OSMAN (Open Source Metric for Measuring Arabic Narratives) - a novel open source Arabic readability metric and tool. It allows researchers to calculate readability for Arabic text with and without diacritics. …

Strategies for Arabic Readability Modeling

2024-07-03 · Juan Piñeros Liberato, Bashar Alhafni, Muhamed Al Khalil, Nizar Habash

Automatic readability assessment is relevant to building NLP applications for education, content analysis, and accessibility. However, Arabic readability assessment is a challenging task due to Arabic's morphological ric…

Sentence

An Online Readability Leveled Arabic Thesaurus

2020-12-01 · COLING 2020 8 · Zhengyang Jiang, Nizar Habash, Muhamed Al Khalil

This demo paper introduces the online Readability Leveled Arabic Thesaurus interface. For a given user input word, this interface provides the word{'}s possible lemmas, roots, English glosses, related Arabic words and ph…

English to Arabic machine translation of mathematical documents

2023-12-02 · Mustapha Eddahibi, Mohammed Mensouri

This paper is about the development of a machine translation system tailored specifically for LATEX mathematical documents. The system focuses on translating English LATEX mathematical documents into Arabic LATEX, cateri…

Machine TranslationTranslation

Can LLMs Control Readability? A Multi-Dimensional Evaluation Framework for CEFR-Controlled Arabic Generation

2026-06-20 · Nour Rabih, Chatrine Qwaider, Ted Briscoe arxiv

While Large Language Models (LLMs) can generate fluent Arabic text, their ability to reliably control readability levels remains unclear. We propose a multi-dimensional evaluation framework for Common European Framework …

Text Generation