paper-with-me

홈 › Papers

The Interplay of Variant, Size, and Task Type in Arabic Pre-trained Language Models

2021-03-11 · EACL (WANLP) 2021 4 · Go Inoue, Bashar Alhafni, Nurpeiis Baimukan, Houda Bouamor, Nizar Habash

In this paper, we explore the effects of language variants, data sizes, and fine-tuning task types in Arabic pre-trained language models. To do so, we build three pre-trained language models across three variants of Arabic: Modern Standard Arabic (MSA), dialectal Arabic, and classical Arabic, in addition to a fourth language model which is pre-trained on a mix of the three. We also examine the importance of pre-training data size by building additional models that are pre-trained on a scaled-down set of the MSA variant. We compare our different models to each other, as well as to eight publicly available models by fine-tuning them on five NLP tasks spanning 12 datasets. Our results suggest that the variant proximity of pre-training data to fine-tuning data is more important than the pre-training data size. We exploit this insight in defining an optimized system selection model for the studied tasks.

📄 PDF Abstract BibTeX arXiv:2103.06678

Code (1)

CAMeL-Lab/CAMeLBERT 공식 구현 pytorch

Tasks

Language ModelingLanguage Modelling

Similar Papers 제목 키워드 기반

Morphological Analysis and Disambiguation for Gulf Arabic: The Interplay between Resources and Methods

2020-05-01 · LREC 2020 5 · Salam Khalifa, Nasser Zalmout, Nizar Habash

In this paper we present the first full morphological analysis and disambiguation system for Gulf Arabic. We use an existing state-of-the-art morphological disambiguation system to investigate the effects of different da…

Morphological AnalysisMorphological DisambiguationMorphological Tagging

Exploring Tokenization Strategies and Vocabulary Sizes for Enhanced Arabic Language Models

2024-03-17 · Mohamed Taher Alrefaie, Nour Eldin Morsy, Nada Samir

This paper presents a comprehensive examination of the impact of tokenization strategies and vocabulary sizes on the performance of Arabic language models in downstream natural language processing tasks. Our investigatio…

Computational EfficiencyHate Speech DetectionMorphological AnalysisNatural Language Inference+2

Arabic Dialects Identification for All Arabic countries

2020-12-01 · COLING (WANLP) 2020 12 · Ahmed Aliwy, Hawraa Taher, Zena AboAltaheen

Arabic dialects are among of three main variant of Arabic language (Classical Arabic, modern standard Arabic and dialectal Arabic). It has many variants according to the country, city (provinces) or town. In this paper, …

AllClustering

SAMER: A Semi-Automatically Created Lexical Resource for Arabic Verbal Multiword Expressions Tokens Paradigm and their Morphosyntactic Features

2016-12-01 · WS 2016 12 · Mohamed Al-Badrashiny, Abdelati Hawwari, Mahmoud Ghoneim, Mona Diab

Although MWE are relatively morphologically and syntactically fixed expressions, several types of flexibility can be observed in MWE, verbal MWE in particular. Identifying the degree of morphological and syntactic flexib…

Machine TranslationPOS

Tibyan Corpus: Balanced and Comprehensive Error Coverage Corpus Using ChatGPT for Arabic Grammatical Error Correction

2024-11-07 · Ahlam Alrehili, Areej Alhothali

Natural language processing (NLP) utilizes text data augmentation to overcome sample size constraints. Increasing the sample size is a natural and widely used strategy for alleviating these challenges. In this study, we …

Data AugmentationGrammatical Error Correction