paper-with-me

홈 › Papers

Letz Translate: Low-Resource Machine Translation for Luxembourgish

2023-03-02 · Yewei Song, Saad Ezzini, Jacques Klein, Tegawende Bissyande, Clément Lefebvre, Anne Goujon

Natural language processing of Low-Resource Languages (LRL) is often challenged by the lack of data. Therefore, achieving accurate machine translation (MT) in a low-resource environment is a real problem that requires practical solutions. Research in multilingual models have shown that some LRLs can be handled with such models. However, their large size and computational needs make their use in constrained environments (e.g., mobile/IoT devices or limited/old servers) impractical. In this paper, we address this problem by leveraging the power of large multilingual MT models using knowledge distillation. Knowledge distillation can transfer knowledge from a large and complex teacher model to a simpler and smaller student model without losing much in performance. We also make use of high-resource languages that are related or share the same linguistic root as the target LRL. For our evaluation, we consider Luxembourgish as the LRL that shares some roots and properties with German. We build multiple resource-efficient models based on German, knowledge distillation from the multilingual No Language Left Behind (NLLB) model, and pseudo-translation. We find that our efficient models are more than 30\% faster and perform only 4\% lower compared to the large state-of-the-art NLLB model.

📄 PDF Abstract BibTeX arXiv:2303.01347

Code (0)

등록된 구현이 없습니다.

Tasks

Knowledge DistillationMachine TranslationTranslation

Methods 이 논문이 사용한 방법론

Knowledge Distillation A very simple way to improve the performance of almost any machine learning algorithm is to train many different models on the same data and then to average their predictions.…

Similar Papers 제목 키워드 기반

LuxInstruct: A Cross-Lingual Instruction Tuning Dataset For Luxembourgish

2025-10-08 · Fred Philippy, Laura Bernardy, Siwen Guo, Jacques Klein 외 arxiv

Instruction tuning has become a key technique for enhancing the performance of large language models, enabling them to better follow human prompts. However, low-resource languages such as Luxembourgish face severe limita…

Machine Translation

LuxMT Technical Report

2026-02-17 · Nils Rehlinger arxiv

We introduce LuxMT, a machine translation system based on Gemma 3 27B and fine-tuned for translation from Luxembourgish (LB) into French (FR) and English (EN). To assess translation performance, we construct a novel benc…

Machine Translation

Component Analysis of Adjectives in Luxembourgish for Detecting Sentiments

2020-05-01 · LREC 2020 5 · Joshgun Sirajzade, Daniela Gierschek, Christoph Schommer

The aim of this paper is to investigate the role of Luxembourgish adjectives in expressing sentiments in user comments written at the web presence of rtl.lu (RTL is the abbreviation for Radio Television Letzebuerg). Alon…

Sentence

Adapting Multilingual Embedding Models to Historical Luxembourgish

2025-02-11 · Andrianos Michail, Corina Julia Raclé, Juri Opitz, Simon Clematide

The growing volume of digitized historical texts requires effective semantic search using text embeddings. However, pre-trained multilingual models, typically evaluated on contemporary texts, face challenges with histori…

ArticlesOptical Character Recognition (OCR)

LuxemBERT: Simple and Practical Data Augmentation in Language Model Pre-Training for Luxembourgish

2022-06-01 · LREC 2022 6 · Cedric Lothritz, Bertrand Lebichot, Kevin Allix, Lisa Veiber 외

Pre-trained Language Models such as BERT have become ubiquitous in NLP where they have achieved state-of-the-art performance in most NLP tasks. While these models are readily available for English and other widely spoken…

Data AugmentationLanguage ModelingLanguage Modelling