paper-with-me

홈 › Papers

AraMUS: Pushing the Limits of Data and Model Scale for Arabic Natural Language Processing

2023-06-11 · Asaad Alghamdi, Xinyu Duan, Wei Jiang, Zhenhai Wang, Yimeng Wu, Qingrong Xia, Zhefeng Wang, Yi Zheng, Mehdi Rezagholizadeh, Baoxing Huai, Peilun Cheng, Abbas Ghaddar

Developing monolingual large Pre-trained Language Models (PLMs) is shown to be very successful in handling different tasks in Natural Language Processing (NLP). In this work, we present AraMUS, the largest Arabic PLM with 11B parameters trained on 529GB of high-quality Arabic textual data. AraMUS achieves state-of-the-art performances on a diverse set of Arabic classification and generative tasks. Moreover, AraMUS shows impressive few-shot learning abilities compared with the best existing Arabic PLMs.

📄 PDF Abstract BibTeX arXiv:2306.06800

Code (0)

등록된 구현이 없습니다.

Tasks

Few-Shot Learning

Similar Papers 제목 키워드 기반

How Well Do LLMs Understand Tunisian Arabic?

2025-11-12 · Mohamed Mahdi arxiv

Large Language Models (LLMs) are the engines driving today's AI agents. The better these models understand human languages, the more natural and user-friendly the interaction with AI becomes, from everyday devices like c…

Sentiment Analysis

Munsit at NADI 2025 Shared Task 2: Pushing the Boundaries of Multidialectal Arabic ASR with Weakly Supervised Pretraining and Continual Supervised Fine-tuning

2025-08-12 · Mahmoud Salhab, Shameed Sait, Mohammad Abusheikh, Hasan Abusheikh arxiv

Automatic speech recognition (ASR) plays a vital role in enabling natural human-machine interaction across applications such as virtual assistants, industrial automation, customer support, and real-time transcription. Ho…

Speech Recognition

Conceptual 12M: Pushing Web-Scale Image-Text Pre-Training To Recognize Long-Tail Visual Concepts

2021-02-17 · CVPR 2021 1 · Soravit Changpinyo, Piyush Sharma, Nan Ding, Radu Soricut

The availability of large-scale image captioning and visual question answering datasets has contributed significantly to recent successes in vision-and-language pre-training. However, these datasets are often collected w…

Caption GenerationDiversityImage CaptioningQuestion Answering+2

Lexical Induction of Morphological and Orthographic Forms for Low-Resourced Languages

2020-12-01 · MSR (COLING) 2020 12 · Taha Tobaili

In this work we address the issue of high-degree lexical sparsity for non-standard languages under severe circumstance of small resources that are considered insufficient to train recent powerful language models. We prop…

Word Embeddings

A review of sentiment analysis research in Arabic language

2020-05-25 · Oumaima Oueslati, Erik Cambria, Moez Ben HajHmida, Habib OUNELLI

Sentiment analysis is a task of natural language processing which has recently attracted increasing attention. However, sentiment analysis research has mainly been carried out for the English language. Although Arabic is…

Arabic Sentiment AnalysisMachine TranslationSentiment AnalysisTransfer Learning+1