Neural Morphology Dataset and Models for Multiple Languages, from the Large to the Endangered
We train neural models for morphological analysis, generation and lemmatization for morphologically rich languages. We present a method for automatically extracting substantially large amount of training data from FSTs for 22 languages, out of which 17 are endangered. The neural models follow the same tagset as the FSTs in order to make it possible to use them as fallback systems together with the FSTs. The source code, models and datasets have been released on Zenodo.
Code (1)
Tasks
LemmatizationMorphological AnalysisSimilar Papers 제목 키워드 기반
Open-Source Morphology for Endangered Mordvinic Languages
This document describes shared development of finite-state description of two closely related but endangered minority languages, Erzya and Moksha. It touches upon morpholexical unity and diversity of the two languages an…
DiversityUnityFST Morphology for the Endangered Skolt Sami Language
We present advances in the development of a FST-based morphological analyzer and generator for Skolt Sami. Like other minority Uralic languages, Skolt Sami exhibits a rich morphology, on the one hand, and there is little…
Morphological AnalysisUniMorph 4.0: Universal Morphology
The Universal Morphology (UniMorph) project is a collaborative effort providing broad-coverage instantiated normalized morphological inflection tables for hundreds of diverse world languages. The project comprises two ma…
Morphological InflectionMorphological Processing of Low-Resource Languages: Where We Are and What's Next
Automatic morphological processing can aid downstream natural language processing applications, especially for low-resource languages, and assist language documentation efforts for endangered languages. Having long been …
Morphological Processing of Low-Resource Languages: Where We Are and What’s Next
Automatic morphological processing can aid downstream natural language processing applications, especially for low-resource languages, and assist language documentation efforts for endangered languages. Having long been …