paper-with-me

Papers

Guidelines for Fine-grained Sentence-level Arabic Readability Annotation

2024-10-11 · Nizar Habash, Hanada Taha-Thomure, Khalid N. Elmadani, Zeina Zeino, Abdallah Abushmaes

This paper presents the foundational framework and initial findings of the Balanced Arabic Readability Evaluation Corpus (BAREC) project, designed to address the need for comprehensive Arabic language resources aligned with diverse readability levels. Inspired by the Taha/Arabi21 readability reference, BAREC aims to provide a standardized reference for assessing sentence-level Arabic text readability across 19 distinct levels, ranging in targets from kindergarten to postgraduate comprehension. Our ultimate goal with BAREC is to create a comprehensive and balanced corpus that represents a wide range of genres, topics, and regional variations through a multifaceted approach combining manual annotation with AI-driven tools. This paper focuses on our meticulous annotation guidelines, demonstrated through the analysis of 10,631 sentences/phrases (113,651 words). The average pairwise inter-annotator agreement, measured by Quadratic Weighted Kappa, is 79.9%, reflecting a high level of substantial agreement. We also report competitive results for benchmarking automatic readability assessment. We will make the BAREC corpus and guidelines openly accessible to support Arabic language research and education.

📄 PDF Abstract BibTeX arXiv:2410.08674

Code (0)

등록된 구현이 없습니다.

Tasks

BenchmarkingSentence

Similar Papers 제목 키워드 기반

A Large and Balanced Corpus for Fine-grained Arabic Readability Assessment

2025-02-19 · Khalid N. Elmadani, Nizar Habash, Hanada Taha-Thomure

This paper introduces the Balanced Arabic Readability Evaluation Corpus BAREC, a large-scale, fine-grained dataset for Arabic readability assessment. BAREC consists of 68,182 sentences spanning 1+ million words, carefull…

Diversity

Simple But Not Na\"\ive: Fine-Grained Arabic Dialect Identification Using Only N-Grams

2019-08-01 · WS 2019 8 · Sohaila Eltanbouly, May Bashendy, Tamer Elsayed

This paper presents the participation of Qatar University team in MADAR shared task, which addresses the problem of sentence-level fine-grained Arabic Dialect Identification over 25 different Arabic dialects in addition …

Dialect IdentificationSentence

Arabic Fine-Grained Entity Recognition

2023-10-26 · Haneen Liqreina, Mustafa Jarrar, Mohammed Khalilia, Ahmed Oumar El-Shangiti 외

Traditional NER systems are typically trained to recognize coarse-grained entities, and less attention is given to classifying entities into a hierarchy of fine-grained lower-level subtypes. This article aims to advance …

NER

AWATIF: A Multi-Genre Corpus for Modern Standard Arabic Subjectivity and Sentiment Analysis

2012-05-01 · LREC 2012 5 · Muhammad Abdul-Mageed, Mona Diab

We present AWATIF, a multi-genre corpus of Modern Standard Arabic (MSA) labeled for subjectivity and sentiment analysis (SSA) at the sentence level. The corpus is labeled using both regular as well as crowd sourcing meth…

Opinion MiningSentenceSentiment Analysis

ArbDialectID at MADAR Shared Task 1: Language Modelling and Ensemble Learning for Fine Grained Arabic Dialect Identification

2019-08-01 · WS 2019 8 · Kathrein Abu Kwaik, Motaz Saad

In this paper, we present a Dialect Identification system (ArbDialectID) that competed at Task 1 of the MADAR shared task, MADARTravel Domain Dialect Identification. We build a course and a fine-grained identification mo…

Dialect IdentificationEnsemble LearningFeature EngineeringLanguage Modelling+1