paper-with-me

Papers

A Multi-Orthography Parallel Corpus of Yiddish Nouns

2020-05-01 · LREC 2020 5 · Jonne Saleva

Yiddish is a low-resource language belonging to the Germanic language family and written using the Hebrew alphabet. As a language, Yiddish can be considered resource-poor as it lacks both public accessible corpora and a widely-used standard orthography, with various countries and organizations influencing the spellings speakers use. While existing corpora of Yiddish text do exist, they are often only written in a single, potentially non-standard orthography, with no parallel version with standard orthography available. In this work, we introduce the first multi-orthography parallel corpus of Yiddish nouns built by scraping word entries from Wiktionary. We also demonstrate how the corpus can be used to bootstrap a transliteration model using the Sequitur-G2P grapheme-to-phoneme conversion toolkit to map between various orthographies. Our trained system achieves error rates between 16.79{\%} and 28.47{\%} on the test set, depending on the orthographies considered. In addition to quantitative analysis, we also conduct qualitative error analysis of the trained system, concluding that non-phonetically spelled Hebrew words are the largest cause of error. We conclude with remarks regarding future work and release the corpus and associated code under a permissive license for the larger community to use.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Tasks

Grapheme-to-Phoneme ConversionTransliteration

Similar Papers 제목 키워드 기반

A Part-of-Speech Tagger for Yiddish

2022-04-03 · Seth Kulick, Neville Ryant, Beatrice Santorini, Joel Wallenberg 외

We describe the construction and evaluation of a part-of-speech tagger for Yiddish. This is the first step in a larger project of automatically assigning part-of-speech tags and syntactic structure to Yiddish text for pu…

Word Embeddings

Jochre 3 and the Yiddish OCR corpus

2025-01-14 · Assaf Urieli, Amber Clooney, Michelle Sigiel, Grisha Leyfer

We describe the construction of a publicly available Yiddish OCR Corpus, and describe and evaluate the open source OCR tool suite Jochre 3, including an Alto editor for corpus annotation, OCR software for Alto OCR layer …

Optical Character Recognition (OCR)

Generating a Yiddish Speech Corpus, Forced Aligner and Basic ASR System for the AHEYM Project

2016-05-01 · LREC 2016 5 · Malgorzata {\'C}avar, Damir {\'C}avar, Dov-Ber Kerler, Anya Quilitzsch

To create automatic transcription and annotation tools for the AHEYM corpus of recorded interviews with Yiddish speakers in Eastern Europe we develop initial Yiddish language resources that are used for adaptations of sp…

The First Parallel Multilingual Corpus of Persian: Toward a Persian BLARK

2014-04-17 · Behrang Qasemizadeh, Saeed Rahimi, Behrooz Mahmoodi Bakhtiari

In this article, we have introduced the first parallel corpus of Persian with more than 10 other European languages. This article describes primary steps toward preparing a Basic Language Resources Kit (BLARK) for Persia…

MameLoshnLM: Yiddish Language Model and Evaluation Benchmark

2026-08-06 · Uri Katz, Omer Goldman, Tomasz Limisiewicz, Reut Tsarfaty 외 arxiv

We present MameLoshnLM, the first open-source 8B-parameter language model built specifically for Yiddish. Despite Yiddish's rich textual tradition, its limited digital presence and the scarcity of reliable evaluation res…

Information Extraction