paper-with-me

홈 › Papers

Nile-Chat: Egyptian Language Models for Arabic and Latin Scripts

2025-07-06 · Guokan Shang, Hadi Abdine, Ahmad Chamma, Amr Mohamed, Mohamed Anwar, Abdelaziz Bounhar, Omar El Herraoui, Preslav Nakov, Michalis Vazirgiannis, Eric Xing arxiv

We introduce Nile-Chat-4B, 3x4B-A6B, and 12B, a collection of LLMs for Egyptian dialect, uniquely designed to understand and generate texts written in both Arabic and Latin scripts. Specifically, with Nile-Chat-3x4B-A6B, we introduce a novel language adaptation approach by leveraging the Branch-Train-MiX strategy to merge script-specialized experts, into a single MoE model. Our Nile-Chat models significantly outperform leading multilingual and Arabic LLMs, such as LLaMa, Jais, and ALLaM, on our newly introduced Egyptian evaluation benchmarks, which span both understanding and generative tasks. Notably, our 12B model yields a 14.4% performance gain over Qwen2.5-14B-Instruct on Latin-script benchmarks. All our resources are publicly available. We believe this work presents a comprehensive methodology for adapting LLMs to dual-script languages, addressing an often overlooked aspect in modern LLM development.

📄 PDF Abstract BibTeX arXiv:2507.04569

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

NileChat: Towards Linguistically Diverse and Culturally Aware LLMs for Local Communities

2025-05-23 · Abdellah El Mekki, Houdaifa Atou, Omer Nacar, Shady Shehata 외

Enhancing the linguistic capabilities of Large Language Models (LLMs) to include low-resource languages is a critical research area. Current research directions predominantly rely on synthetic data generated by translati…

Translation

LLM-to-Speech: A Synthetic Data Pipeline for Training Dialectal Text-to-Speech Models

2026-02-17 · Ahmed Khaled Khamis, Hesham Ali arxiv

Despite the advances in neural text to speech (TTS), many Arabic dialectal varieties remain marginally addressed, with most resources concentrated on Modern Spoken Arabic (MSA) and Gulf dialects, leaving Egyptian Arabic …

Synthetic Data GenerationSpeaker DiarizationSpeech SynthesisText to Speech

NileULex: A Phrase and Word Level Sentiment Lexicon for Egyptian and Modern Standard Arabic

2016-05-01 · LREC 2016 5 · Samhaa R. El-Beltagy

This paper presents NileULex, which is an Arabic sentiment lexicon containing close to six thousands Arabic words and compound phrases. Forty five percent of the terms and expressions in the lexicon are Egyptian or collo…

Sentiment AnalysisTranslation

Botta: An Arabic Dialect Chatbot

2016-12-01 · COLING 2016 12 · Dana Abu Ali, Nizar Habash

This paper presents BOTTA, the first Arabic dialect chatbot. We explore the challenges of creating a conversational agent that aims to simulate friendly conversations using the Egyptian Arabic dialect. We present a numbe…

Chatbot

Egyptian Arabic to English Statistical Machine Translation System for NIST OpenMT'2015

2016-06-18 · Hassan Sajjad, Nadir Durrani, Francisco Guzman, Preslav Nakov 외

The paper describes the Egyptian Arabic-to-English statistical machine translation (SMT) system that the QCRI-Columbia-NYUAD (QCN) group submitted to the NIST OpenMT'2015 competition. The competition focused on informal …

Language ModelingLanguage ModellingMachine TranslationTranslation+1