paper-with-me

홈 › Papers

LuxIT: A Luxembourgish Instruction Tuning Dataset from Monolingual Seed Data

2025-10-28 · Julian Valline, Cedric Lothritz, Siwen Guo, Jordi Cabot arxiv

The effectiveness of instruction-tuned Large Language Models (LLMs) is often limited in low-resource linguistic settings due to a lack of high-quality training data. We introduce LuxIT, a novel, monolingual instruction tuning dataset for Luxembourgish developed to mitigate this challenge. We synthesize the dataset from a corpus of native Luxembourgish texts, utilizing DeepSeek-R1-0528, chosen for its shown proficiency in Luxembourgish. Following generation, we apply a quality assurance process, employing an LLM-as-a-judge approach, retaining 227,507 high-quality instruction-answer pairs. To investigate the practical utility of the dataset, we fine-tune 14 smaller-scale LLMs ($\leq$15B parameters) on LuxIT and evaluate them on standardized Luxembourgish proficiency exams and five downstream NLP tasks. Training on LuxIT yields a mean accuracy change of +5.37 percentage points on language exams across all 14 models, with 12 of 14 showing improvement. On NLP downstream tasks, 9 of 14 models improve in macro-averaged F1, though gains on the two benchmarks do not systematically correlate. These results underscore the feasibility of leveraging monolingual synthetic data to improve LLM capabilities in low-resource languages, while highlighting the multi-faceted nature of language proficiency.

📄 PDF Abstract BibTeX arXiv:2510.24434

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

LuxInstruct: A Cross-Lingual Instruction Tuning Dataset For Luxembourgish

2025-10-08 · Fred Philippy, Laura Bernardy, Siwen Guo, Jacques Klein 외 arxiv

Instruction tuning has become a key technique for enhancing the performance of large language models, enabling them to better follow human prompts. However, low-resource languages such as Luxembourgish face severe limita…

Machine Translation

Text Generation Models for Luxembourgish with Limited Data: A Balanced Multilingual Strategy

2024-12-12 · Alistair Plum, Tharindu Ranasinghe, Christoph Purschke

This paper addresses the challenges in developing language models for less-represented languages, with a focus on Luxembourgish. Despite its active development, Luxembourgish faces a digital data scarcity, exacerbated by…

Cross-Lingual TransferText GenerationTransfer Learning

LuxEmbedder: A Cross-Lingual Approach to Enhanced Luxembourgish Sentence Embeddings

2024-12-04 · Fred Philippy, Siwen Guo, Jacques Klein, Tegawendé F. Bissyandé

Sentence embedding models play a key role in various Natural Language Processing tasks, such as in Topic Modeling, Document Clustering and Recommendation Systems. However, these models rely heavily on parallel data, whic…

Recommendation SystemsSentenceSentence EmbeddingSentence-Embedding+1

Multilingual Instruction Tuning With Just a Pinch of Multilinguality

2024-01-03 · Uri Shaham, Jonathan Herzig, Roee Aharoni, Idan Szpektor 외

As instruction-tuned large language models (LLMs) gain global adoption, their ability to follow instructions in multiple languages becomes increasingly crucial. In this work, we investigate how multilinguality during ins…

Cross-Lingual TransferInstruction Following

Teaching Llama a New Language Through Cross-Lingual Knowledge Transfer

2024-04-05 · Hele-Andra Kuulmets, Taido Purason, Agnes Luhtaru, Mark Fishel

This paper explores cost-efficient methods to adapt pretrained Large Language Models (LLMs) to new lower-resource languages, with a specific focus on Estonian. Leveraging the Llama 2 model, we investigate the impact of c…

Instruction FollowingTransfer Learning