paper-with-me

홈 › Papers

The Role of Orthographic Consistency in Multilingual Embedding Models for Text Classification in Arabic-Script Languages

2025-07-24 · Abdulhady Abas Abdullah, Amir H. Gandomi, Tarik A Rashid, Seyedali Mirjalili, Laith Abualigah, Milena Živković, Hadi Veisi arxiv

In natural language processing, multilingual models like mBERT and XLM-RoBERTa promise broad coverage but often struggle with languages that share a script yet differ in orthographic norms and cultural context. This issue is especially notable in Arabic-script languages such as Kurdish Sorani, Arabic, Persian, and Urdu. We introduce the Arabic Script RoBERTa (AS-RoBERTa) family: four RoBERTa-based models, each pre-trained on a large corpus tailored to its specific language. By focusing pre-training on language-specific script features and statistics, our models capture patterns overlooked by general-purpose models. When fine-tuned on classification tasks, AS-RoBERTa variants outperform mBERT and XLM-RoBERTa by 2 to 5 percentage points. An ablation study confirms that script-focused pre-training is central to these gains. Error analysis using confusion matrices shows how shared script traits and domain-specific content affect performance. Our results highlight the value of script-aware specialization for languages using the Arabic script and support further work on pre-training strategies rooted in script and language specificity.

📄 PDF Abstract BibTeX arXiv:2507.18762

Code (0)

등록된 구현이 없습니다.

Tasks

Text Classification

Similar Papers 제목 키워드 기반

Enhancing Robustness of Autoregressive Language Models against Orthographic Attacks via Pixel-based Approach

2025-08-28 · Han Yang, Jian Lan, Yihong Liu, Hinrich Schütze 외 arxiv

Autoregressive language models are vulnerable to orthographic attacks, where input text is perturbed with characters from multilingual alphabets, leading to substantial performance degradation. This vulnerability primari…

Data-adaptive Transfer Learning for Translation: A Case Study in Haitian and Jamaican

2022-09-13 · loresmt (COLING) 2022 10 · Nathaniel R. Robinson, Cameron J. Hogan, Nancy Fulda, David R. Mortensen

Multilingual transfer techniques often improve low-resource machine translation (MT). Many of these techniques are applied without considering data characteristics. We show in the context of Haitian-to-English translatio…

Cross-Lingual TransferMachine TranslationTransfer LearningTranslation

Data-adaptive Transfer Learning for Low-resource Translation: A Case Study in Haitian

2021-11-16 · ACL ARR November 2021 11 · Anonymous

Multilingual transfer techniques often improve low-resource machine translation (MT). Many of these techniques are applied without considering data characteristics. We show in the context of Haitian-to-English translatio…

Cross-Lingual TransferMachine TranslationTransfer LearningTranslation

Not Truly Multilingual: Script Consistency as a Missing Dimension in VLM Evaluation

2026-06-15 · Prabhjot Singh, Bhushan Pawar, Madhu Reddiboina, Rajvee Sheth arxiv

Current multilingual evaluations for Vision-Language Models (VLMs) assume a one-to-one mapping between language and orthography, overlooking billions of users of multi-script languages. We introduce PuMVR (Punjabi Multim…

Visual Reasoning

Machine Translation by Projecting Text into the Same Phonetic-Orthographic Space Using a Common Encoding

2023-05-21 · Amit Kumar, Shantipriya Parida, Ajay Pratap, Anil Kumar Singh

The use of subword embedding has proved to be a major innovation in Neural Machine Translation (NMT). It helps NMT to learn better context vectors for Low Resource Languages (LRLs) so as to predict the target words by be…

Machine TranslationNMTTranslation