paper-with-me

Papers

LLiMba: Sardinian on a Single GPU -- Adapting a 3B Language Model to a Vanishing Romance Language

2026-05-09 · Luca Ballore arxiv

Sardinian, a Romance language with roughly one million speakers, has minimal presence in modern NLP. Commercial services do not support it, and current language models do not produce it reliably. We present LLiMba, a 3B parameter Sardinian-ready model adapted from Qwen2.5-3B-Instruct through continued pretraining (CPT) and supervised fine-tuning (SFT) on a single 24 GB consumer GPU. The corpus contains 11.5 million tokens of Sardinian spanning LSC, Logudorese, and Campidanese, augmented with 2.4 million tokens of related Romance text as replay against register blurring. After CPT the model reaches a perplexity of 6.76 on held out Sardinian and outperforms the base across all six FLORES-200 directions. We compare five SFT configurations under matched conditions: full fine-tuning, LoRA r64, rsLoRA r128, rsLoRA r256, and DoRA r256. rsLoRA r256 wins on every direction into Sardinian, reaching 28.5 BLEU from English against 17.3 after CPT and 21.0 with full fine-tuning. The rank ablation places r128 between LoRA r64 and rsLoRA r256 on BLEU but reveals failure modes invisible to the metric, including leakage across scripts no other variant produces. LoRA r64 retains less factual content from SFT than configurations at higher rank and produces more confident fabrications, though all methods fabricate on content absent from training. DoRA r256 yields the smallest gap between training and evaluation but the worst factual accuracy. The findings indicate that adapter capacity matters more than the choice among LoRA variants for adapting a Romance pretrained base to a low resource Romance target, that stronger regularization is not uniformly beneficial, and that translation metrics smoothly order configurations whose qualitative behavior differs categorically. Perplexity comparisons across scripts must account for byte fallback tokenization, which deflates the metric for scripts other than Latin.

📄 PDF Abstract BibTeX arXiv:2605.09015

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

SardaNet: a Linguistic Resource for Sardinian Language

2018-01-01 · GWC 2018 1 · Manuela Angioni, Franco Tuveri, Maurizio Virdis, Laura Lucia Lai 외

This paper describes the process of building SardaNet, a linguistic resource for Sardinian language including the different linguistic varieties in Sardinia. SardaNet aims at identifying the semantic relations between Sa…

A Bolu: A Structured Dataset for the Computational Analysis of Sardinian Improvisational Poetry

2026-04-21 · Silvio Calderaro, Johanna Monti arxiv

The growing interest of Natural Language Processing (NLP) in minority languages has not yet bridged the gap in the preservation of oral linguistic heritage. In particular, extemporaneous poetry - a performative genre bas…

When Less Is More? Diagnosing ASR Predictions in Sardinian via Layer-Wise Decoding

2026-02-10 · Domenico De Cristofaro, Alessandro Vietti, Marianne Pouplier, Aleese Block arxiv

Recent studies have shown that intermediate layers in multilingual speech models often encode more phonetically accurate representations than the final output layer. In this work, we apply a layer-wise decoding strategy …

Initial Experiments In Cross-Lingual Morphological Analysis Using Morpheme Segmentation

2019-06-01 · WS 2019 6 · Vladislav Mikhailov, Lorenzo Tosi, Anastasia Khorosheva, Oleg Serikov

The paper describes initial experiments in data-driven cross-lingual morphological analysis of open-category words using a combination of unsupervised morpheme segmentation, annotation projection and an LSTM encoder-deco…

DecoderMorphological Analysis

LIMBA: An Open-Source Framework for the Preservation and Valorization of Low-Resource Languages using Generative Models

2024-11-20 · Salvatore Mario Carta, Stefano Chessa, Giulia Contu, Andrea Corriga 외

Minority languages are vital to preserving cultural heritage, yet they face growing risks of extinction due to limited digital resources and the dominance of artificial intelligence models trained on high-resource langua…

Diversity