paper-with-me

Papers

A Large and Balanced Corpus for Fine-grained Arabic Readability Assessment

2025-02-19 · Khalid N. Elmadani, Nizar Habash, Hanada Taha-Thomure

This paper introduces the Balanced Arabic Readability Evaluation Corpus BAREC, a large-scale, fine-grained dataset for Arabic readability assessment. BAREC consists of 68,182 sentences spanning 1+ million words, carefully curated to cover 19 readability levels, from kindergarten to postgraduate comprehension. The corpus balances genre diversity, topical coverage, and target audiences, offering a comprehensive resource for evaluating Arabic text complexity. The corpus was fully manually annotated by a large team of annotators. The average pairwise inter-annotator agreement, measured by Quadratic Weighted Kappa, is 81.3%, reflecting a high level of substantial agreement. Beyond presenting the corpus, we benchmark automatic readability assessment across different granularity levels, comparing a range of techniques. Our results highlight the challenges and opportunities in Arabic readability modeling, demonstrating competitive performance across various methods. To support research and education, we will make BAREC openly available, along with detailed annotation guidelines and benchmark results.

📄 PDF Abstract BibTeX arXiv:2502.13520

Code (0)

등록된 구현이 없습니다.

Tasks

Diversity

Similar Papers 제목 키워드 기반

Guidelines for Fine-grained Sentence-level Arabic Readability Annotation

2024-10-11 · Nizar Habash, Hanada Taha-Thomure, Khalid N. Elmadani, Zeina Zeino 외

This paper presents the foundational framework and initial findings of the Balanced Arabic Readability Evaluation Corpus (BAREC) project, designed to address the need for comprehensive Arabic language resources aligned w…

BenchmarkingSentence

A Fine-Grained Annotated Multi-Dialectal Arabic Corpus

2019-09-01 · RANLP 2019 9 · Anis Charfi, Wajdi Zaghouani, Syed Hassan Mehdi, Esraa Mohamed

We present ARAP-Tweet 2.0, a corpus of 5 million dialectal Arabic tweets and 50 million words of about 3000 Twitter users from 17 Arab countries. Compared to the first version, the new corpus has significant improvements…

ARCADE: A City-Scale Corpus for Fine-Grained Arabic Dialect Tagging

2026-01-05 · Omer Nacar, Serry Sibaee, Adel Ammar, Yasser Alhabashi 외 arxiv

The Arabic language is characterized by a rich tapestry of regional dialects that differ substantially in phonetics and lexicon, reflecting the geographic and cultural diversity of its speakers. Despite the availability …

Multi-Task Learning

Automatically Developing a Fine-grained Arabic Named Entity Corpus and Gazetteer by utilizing Wikipedia

2013-10-01 · IJCNLP 2013 10 · Fahd Alotaibi, Mark Lee
Question AnsweringTransliteration

Arabic Fine-Grained Entity Recognition

2023-10-26 · Haneen Liqreina, Mustafa Jarrar, Mohammed Khalilia, Ahmed Oumar El-Shangiti 외

Traditional NER systems are typically trained to recognize coarse-grained entities, and less attention is given to classifying entities into a hierarchy of fine-grained lower-level subtypes. This article aims to advance …

NER