paper-with-me

홈 › Papers

!MSA at BAREC Shared Task 2025: Ensembling Arabic Transformers for Readability Assessment

2025-09-12 · Mohamed Basem, Mohamed Younes, Seif Ahmed, Abdelrahman Moustafa arxiv

We present MSAs winning system for the BAREC 2025 Shared Task on fine-grained Arabic readability assessment, achieving first place in six of six tracks. Our approach is a confidence-weighted ensemble of four complementary transformer models (AraBERTv2, AraELECTRA, MARBERT, and CAMeLBERT) each fine-tuned with distinct loss functions to capture diverse readability signals. To tackle severe class imbalance and data scarcity, we applied weighted training, advanced preprocessing, SAMER corpus relabeling with our strongest model, and synthetic data generation via Gemini 2.5 Flash, adding about 10,000 rare-level samples. A targeted post-processing step corrected prediction distribution skew, delivering a 6.3 percent Quadratic Weighted Kappa (QWK) gain. Our system reached 87.5 percent QWK at the sentence level and 87.4 percent at the document level, demonstrating the power of model and loss diversity, confidence-informed fusion, and intelligent augmentation for robust Arabic readability prediction.

📄 PDF Abstract BibTeX arXiv:2509.10040

Code (0)

등록된 구현이 없습니다.

Tasks

Synthetic Data Generation

Similar Papers 제목 키워드 기반

mucAI at BAREC Shared Task 2025: Towards Uncertainty Aware Arabic Readability Assessment

2025-09-18 · Ahmed Abdou arxiv

We present a simple, model-agnostic post-processing technique for fine-grained Arabic readability classification in the BAREC 2025 Shared Task (19 ordinal levels). Our method applies conformal prediction to generate pred…

Guidelines for Fine-grained Sentence-level Arabic Readability Annotation

2024-10-11 · Nizar Habash, Hanada Taha-Thomure, Khalid N. Elmadani, Zeina Zeino 외

This paper presents the foundational framework and initial findings of the Balanced Arabic Readability Evaluation Corpus (BAREC) project, designed to address the need for comprehensive Arabic language resources aligned w…

BenchmarkingSentence

A Large and Balanced Corpus for Fine-grained Arabic Readability Assessment

2025-02-19 · Khalid N. Elmadani, Nizar Habash, Hanada Taha-Thomure

This paper introduces the Balanced Arabic Readability Evaluation Corpus BAREC, a large-scale, fine-grained dataset for Arabic readability assessment. BAREC consists of 68,182 sentences spanning 1+ million words, carefull…

Diversity

Lexicon-Enriched Graph Modeling for Arabic Document Readability Prediction

2025-09-26 · Passant Elchafei, Mayar Osama, Mohamed Rageh, Mervat Abuelkheir arxiv

We present a graph-based approach enriched with lexicons to predict document-level readability in Arabic, developed as part of the Constrained Track of the BAREC Shared Task 2025. Our system models each document as a sen…

Graph Neural Network

ImageEval 2026: Culturally Grounded Arabic Multimodal Evaluation

2026-08-31 · Samir Abdaljalil, Hunzalah Hassan Bhatti, Ahlam Bashiti, Farina Amir 외 arxiv

We present an overview of the ImageEval 2026 shared task on culturally grounded Arabic multimodal evaluation. It includes two tasks: (i) AynVQA, covering spoken visual question answering and image-grounded hallucination …

Visual Question AnsweringText-to-Image Generation