paper-with-me

홈 › Papers

Tahakom LLM Guidelines and Recipes: From Pre-training Data to an Arabic LLM

2025-10-15 · Areej AlOtaibi, Lina Alyahya, Raghad Alshabanah, Shahad Alfawzan, Shuruq Alarefei, Reem Alsabti, Nouf Alsubaie, Abdulaziz Alhuzaymi, Lujain Alkhelb, Majd Alsayari, Waad Alahmed, Omar Talabay, Jalal Alowibdi, Salem Alelyani, Adel Bibi arxiv

Large Language Models (LLMs) have significantly advanced the field of natural language processing, enhancing capabilities in both language understanding and generation across diverse domains. However, developing LLMs for Arabic presents unique challenges. This paper explores these challenges by focusing on critical aspects such as data curation, tokenizer design, and evaluation. We detail our approach to the collection and filtration of Arabic pre-training datasets, assess the impact of various tokenizer designs on model performance, and examine the limitations of existing Arabic evaluation frameworks, for which we propose a systematic corrective methodology. To promote transparency and facilitate collaborative development, we share our data and methodologies, contributing to the advancement of language modeling, particularly for the Arabic language.

📄 PDF Abstract BibTeX arXiv:2510.13481

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Large Scale Arabic Error Annotation: Guidelines and Framework

2014-05-01 · LREC 2014 5 · Wajdi Zaghouani, Behrang Mohit, Nizar Habash, Ossama Obeid 외

We present annotation guidelines and a web-based annotation framework developed as part of an effort to create a manually annotated Arabic corpus of errors and corrections for various text types. Such a corpus will be in…

Machine Translation

Guidelines and Framework for a Large Scale Arabic Diacritized Corpus

2016-05-01 · LREC 2016 5 · Wajdi Zaghouani, Houda Bouamor, Abdelati Hawwari, Mona Diab 외

This paper presents the annotation guidelines developed as part of an effort to create a large scale manually diacritized corpus for various Arabic text genres. The target size of the annotated corpus is 2 million words.…

Building an Arabic Machine Translation Post-Edited Corpus: Guidelines and Annotation

2016-05-01 · LREC 2016 5 · Wajdi Zaghouani, Nizar Habash, Ossama Obeid, Behrang Mohit 외

We present our guidelines and annotation procedure to create a human corrected machine translated post-edited corpus for the Modern Standard Arabic. Our overarching goal is to use the annotated corpus to develop automati…

ArticlesMachine TranslationTranslation

Simplified guidelines for the creation of Large Scale Dialectal Arabic Annotations

2012-05-01 · LREC 2012 5 · Heba Elfardy, Mona Diab

The Arabic language is a collection of dialectal variants along with the standard form, Modern Standard Arabic (MSA). MSA is used in official Settings while the dialectal variants (DA) correspond to the native tongue of …

Speech Recognition

Guidelines and Annotation Framework for Arabic Author Profiling

2018-08-23 · Wajdi Zaghouani, Anis Charfi

In this paper, we present the annotation pipeline and the guidelines we wrote as part of an effort to create a large manually annotated Arabic author profiling dataset from various social media sources covering 16 Arabic…

Author Profiling