paper-with-me

Papers

MameLoshnLM: Yiddish Language Model and Evaluation Benchmark

2026-08-06 · Uri Katz, Omer Goldman, Tomasz Limisiewicz, Reut Tsarfaty, Noah A. Smith arxiv

We present MameLoshnLM, the first open-source 8B-parameter language model built specifically for Yiddish. Despite Yiddish's rich textual tradition, its limited digital presence and the scarcity of reliable evaluation resources have constrained progress in Yiddish language modeling. Existing multilingual corpora and benchmarks are often poor proxies for the language, containing substantial amounts of noisy, machine-translated, and misclassified text. We address these gaps by introducing Oytser, a high-quality Yiddish pretraining corpus that combines contemporary web-native sources with literary materials, and Kashes, a multi-task benchmark spanning translation, linguistic analysis, information extraction, and language understanding. Using these resources, we continue pretraining Llama 3.1 8B to obtain MameLoshnLM. Across the tasks in the benchmark, MameLoshnLM outperforms open baselines of similar scale. Our analyses show that these gains are not only quantitative: relative to general-purpose multilingual models, MameLoshnLM better captures language-defining lexical and morphological patterns, pointing to a broader failure mode of noisy web-scale multilingual data for low-resource languages. Our results provide both a foundation for Yiddish NLP and a practical template for language model development in historically rich but digitally underrepresented languages.

📄 PDF Abstract BibTeX arXiv:2608.05850

Code (0)

등록된 구현이 없습니다.

Tasks

Information Extraction

Similar Papers 제목 키워드 기반

A Digital Swedish-Yiddish/Yiddish-Swedish Dictionary: A Web-Based Dictionary that is also Available Offline

2022-06-01 · EURALI (LREC) 2022 6 · Magnus Ahltorp, Jean Hessel, Gunnar Eriksson, Maria Skeppstedt 외

Yiddish is one of the national minority languages of Sweden, and one of the languages for which the Swedish Institute for Language and Folklore is responsible for developing useful language resources. We here describe th…

Transliteration

Generating a Yiddish Speech Corpus, Forced Aligner and Basic ASR System for the AHEYM Project

2016-05-01 · LREC 2016 5 · Malgorzata {\'C}avar, Damir {\'C}avar, Dov-Ber Kerler, Anya Quilitzsch

To create automatic transcription and annotation tools for the AHEYM corpus of recorded interviews with Yiddish speakers in Eastern Europe we develop initial Yiddish language resources that are used for adaptations of sp…

A Part-of-Speech Tagger for Yiddish

2022-04-03 · Seth Kulick, Neville Ryant, Beatrice Santorini, Joel Wallenberg 외

We describe the construction and evaluation of a part-of-speech tagger for Yiddish. This is the first step in a larger project of automatically assigning part-of-speech tags and syntactic structure to Yiddish text for pu…

Word Embeddings

A Multi-Orthography Parallel Corpus of Yiddish Nouns

2020-05-01 · LREC 2020 5 · Jonne Saleva

Yiddish is a low-resource language belonging to the Germanic language family and written using the Hebrew alphabet. As a language, Yiddish can be considered resource-poor as it lacks both public accessible corpora and a …

Grapheme-to-Phoneme ConversionTransliteration

Jochre 3 and the Yiddish OCR corpus

2025-01-14 · Assaf Urieli, Amber Clooney, Michelle Sigiel, Grisha Leyfer

We describe the construction of a publicly available Yiddish OCR Corpus, and describe and evaluate the open source OCR tool suite Jochre 3, including an Alto editor for corpus annotation, OCR software for Alto OCR layer …

Optical Character Recognition (OCR)