paper-with-me

홈 › Papers

chDzDT: Word-level morphology-aware language model for Algerian social media text

2025-09-01 · Abdelkrime Aries arxiv

Pre-trained language models (PLMs) have substantially advanced natural language processing by providing context-sensitive text representations. However, the Algerian dialect remains under-represented, with few dedicated models available. Processing this dialect is challenging due to its complex morphology, frequent code-switching, multiple scripts, and strong lexical influences from other languages. These characteristics complicate tokenization and reduce the effectiveness of conventional word- or subword-level approaches. To address this gap, we introduce chDzDT, a character-level pre-trained language model tailored for Algerian morphology. Unlike conventional PLMs that rely on token sequences, chDzDT is trained on isolated words. This design allows the model to encode morphological patterns robustly, without depending on token boundaries or standardized orthography. The training corpus draws from diverse sources, including YouTube comments, French, English, and Berber Wikipedia, as well as the Tatoeba project. It covers multiple scripts and linguistic varieties, resulting in a substantial pre-training workload. Our contributions are threefold: (i) a detailed morphological analysis of Algerian dialect using YouTube comments; (ii) the construction of a multilingual Algerian lexicon dataset; and (iii) the development and extensive evaluation of a character-level PLM as a morphology-focused encoder for downstream tasks. The proposed approach demonstrates the potential of character-level modeling for morphologically rich, low-resource dialects and lays a foundation for more inclusive and adaptable NLP systems.

📄 PDF Abstract BibTeX arXiv:2509.01772

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Optimal Turkish Subword Strategies at Scale: Systematic Evaluation of Data, Vocabulary, Morphology Interplay

2026-02-06 · Duygu Altinok arxiv

Tokenization is a pivotal design choice for neural language modeling in morphologically rich languages (MRLs) such as Turkish, where productive agglutination challenges both vocabulary efficiency and morphological fideli…

Dependency ParsingSentiment Analysis

Evaluation of Morphological Embeddings for the Russian Language

2021-03-11 · Vitaly Romanov, Albina Khusainova

A number of morphology-based word embedding models were introduced in recent years. However, their evaluation was mostly limited to English, which is known to be a morphologically simple language. In this paper, we explo…

ChunkingNERPOSPOS Tagging+1

Improving Character-Aware Neural Language Model by Warming up Character Encoder under Skip-gram Architecture

2021-09-01 · RANLP 2021 9 · Yukun Feng, Chenlong Hu, Hidetaka Kamigaito, Hiroya Takamura 외

Character-aware neural language models can capture the relationship between words by exploiting character-level information and are particularly effective for languages with rich morphology. However, these models are usu…

Language ModelingLanguage Modelling

Morphologically Aware Word-Level Translation

2020-11-15 · COLING 2020 8 · Paula Czarnowska, Sebastian Ruder, Ryan Cotterell, Ann Copestake

We propose a novel morphologically aware probability model for bilingual lexicon induction, which jointly models lexeme translation and inflectional morphology in a structured way. Our model exploits the basic linguistic…

Bilingual Lexicon InductionTranslation

Morphology Without Borders: Clause-Level Morphology

2022-02-25 · Omer Goldman, Reut Tsarfaty

Morphological tasks use large multi-lingual datasets that organize words into inflection tables, which then serve as training and evaluation data for various tasks. However, a closer inspection of these data reveals prof…