paper-with-me

홈 › Papers

LuxEmbedder: A Cross-Lingual Approach to Enhanced Luxembourgish Sentence Embeddings

2024-12-04 · Fred Philippy, Siwen Guo, Jacques Klein, Tegawendé F. Bissyandé

Sentence embedding models play a key role in various Natural Language Processing tasks, such as in Topic Modeling, Document Clustering and Recommendation Systems. However, these models rely heavily on parallel data, which can be scarce for many low-resource languages, including Luxembourgish. This scarcity results in suboptimal performance of monolingual and cross-lingual sentence embedding models for these languages. To address this issue, we compile a relatively small but high-quality human-generated cross-lingual parallel dataset to train LuxEmbedder, an enhanced sentence embedding model for Luxembourgish with strong cross-lingual capabilities. Additionally, we present evidence suggesting that including low-resource languages in parallel training datasets can be more advantageous for other low-resource languages than relying solely on high-resource language pairs. Furthermore, recognizing the lack of sentence embedding benchmarks for low-resource languages, we create a paraphrase detection benchmark specifically for Luxembourgish, aiming to partially fill this gap and promote further research.

📄 PDF Abstract BibTeX arXiv:2412.03331

Code (1)

fredxlpy/luxembedder 공식 구현

Tasks

Recommendation SystemsSentenceSentence EmbeddingSentence-EmbeddingSentence Embeddings

Similar Papers 제목 키워드 기반

LuxMT Technical Report

2026-02-17 · Nils Rehlinger arxiv

We introduce LuxMT, a machine translation system based on Gemma 3 27B and fine-tuned for translation from Luxembourgish (LB) into French (FR) and English (EN). To assess translation performance, we construct a novel benc…

Machine Translation

Adapting Multilingual Embedding Models to Historical Luxembourgish

2025-02-11 · Andrianos Michail, Corina Julia Raclé, Juri Opitz, Simon Clematide

The growing volume of digitized historical texts requires effective semantic search using text embeddings. However, pre-trained multilingual models, typically evaluated on contemporary texts, face challenges with histori…

ArticlesOptical Character Recognition (OCR)

Automatic language identity tagging on word and sentence-level in multilingual text sources: a case-study on Luxembourgish

2014-05-01 · LREC 2014 5 · Thomas Lavergne, Gilles Adda, Martine Adda-Decker, Lori Lamel

Luxembourgish, embedded in a multilingual context on the divide between Romance and Germanic cultures, remains one of Europe{'}s under-described languages. This is due to the fact that the written production remains rela…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Language IdentificationLanguage Modeling+4

Text Generation Models for Luxembourgish with Limited Data: A Balanced Multilingual Strategy

2024-12-12 · Alistair Plum, Tharindu Ranasinghe, Christoph Purschke

This paper addresses the challenges in developing language models for less-represented languages, with a focus on Luxembourgish. Despite its active development, Luxembourgish faces a digital data scarcity, exacerbated by…

Cross-Lingual TransferText GenerationTransfer Learning

LuxInstruct: A Cross-Lingual Instruction Tuning Dataset For Luxembourgish

2025-10-08 · Fred Philippy, Laura Bernardy, Siwen Guo, Jacques Klein 외 arxiv

Instruction tuning has become a key technique for enhancing the performance of large language models, enabling them to better follow human prompts. However, low-resource languages such as Luxembourgish face severe limita…

Machine Translation