paper-with-me

홈 › Papers

Robust Language Identification for Romansh Varieties

2026-03-16 · Charlotte Model, Sina Ahmadi, Jannis Vamvas arxiv

The Romansh language has several regional varieties, called idioms, which sometimes have limited mutual intelligibility. Despite this linguistic diversity, there has been a lack of documented efforts to build a language identification (LID) system that can distinguish between these idioms. Since Romansh LID should also be able to recognize Rumantsch Grischun, a supra-regional variety that combines elements of several idioms, this makes for a novel and interesting classification problem. In this paper, we present a LID system for Romansh idioms based on an SVM approach. We evaluate our model on a newly curated benchmark across two domains and find that it reaches an average in-domain accuracy of 97%, enabling applications such as idiom-aware spell checking or machine translation. Our classifier is publicly available.

📄 PDF Abstract BibTeX arXiv:2603.15969

Code (0)

등록된 구현이 없습니다.

Tasks

Language IdentificationMachine Translation

Similar Papers 제목 키워드 기반

Translation Asymmetry in LLMs as a Data Augmentation Factor: A Case Study for 6 Romansh Language Varieties

2026-03-26 · Jannis Vamvas, Ignacio Pérez Prat, Angela Heldstab, Dominic P. Fischer 외 arxiv

Recent strategies for low-resource machine translation rely on LLMs to generate synthetic data based on text in higher-resource languages. We revisit this idea for Romansh, a language with 6 distinct varieties. LLMs tend…

Machine TranslationData Augmentation

Expanding the WMT24++ Benchmark with Rumantsch Grischun, Sursilvan, Sutsilvan, Surmiran, Puter, and Vallader

2025-09-03 · Jannis Vamvas, Ignacio Pérez Prat, Not Battesta Soliva, Sandra Baltermia-Guetg 외 arxiv

The Romansh language, spoken in Switzerland, has limited resources for machine translation evaluation. In this paper, we present a benchmark for six varieties of Romansh: Rumantsch Grischun, a supra-regional variety, and…

Machine Translation

The Mediomatix Corpus: Parallel Data for Romansh Language Varieties via Comparable Schoolbooks

2025-08-22 · Zachary Hopton, Jannis Vamvas, Andrin Büchler, Anna Rutkiewicz 외 arxiv

The five idioms (i.e., varieties) of the Romansh language are largely standardized and are taught in the schools of the respective communities in Switzerland. In this paper, we present the first parallel corpus of Romans…

Machine Translation

RUMLEM: A Dictionary-Based Lemmatizer for Romansh

2026-04-13 · Dominic P. Fischer, Zachary Hopton, Jannis Vamvas arxiv

Lemmatization -- the task of mapping an inflected word form to its dictionary form -- is a crucial component of many NLP applications. In this paper, we present RUMLEM, a lemmatizer that covers the five main varieties of…

Does mBERT understand Romansh? Evaluating word embeddings using word alignment

2023-06-14 · Eyal Liron Dolev

We test similarity-based word alignment models (SimAlign and awesome-align) in combination with word embeddings from mBERT and XLM-R on parallel sentences in German and Romansh. Since Romansh is an unseen language, we ar…

SentenceWord AlignmentWord EmbeddingsXLM-R