UralicNLP: An NLP Library for Uralic Languages
UralicNLP is a natural language processing library for small Uralic languages. It can produce morphological analysis, generate morphological forms, lemmatize words and give lexical information about words in Uralic languages. At the time of writing, the following languages are supported: Skolt Sami, Ingrian, Meadow & Eastern Mari, Votic, Olonets-Karelian, Erzya, Moksha, Hill Mari, Udmurt, Tundra Nenets, Komi-Permyak and Finnish. This information originates from FST tools and dictionaries developed in the Giellatekno infrastructure. Currently, UralicNLP uses the nightly builds for languages supported by Apertium and less frequently updated FSTs and CGs for the other languages.
Code (1)
Tasks
Morphological AnalysisSimilar Papers 제목 키워드 기반
Uralic Language Identification (ULI) 2020 shared task dataset and the Wanca 2017 corpus
This article introduces the Wanca 2017 corpus of texts crawled from the internet from which the sentences in rare Uralic languages for the use of the Uralic Language Identification (ULI) 2020 shared task were collected. …
Language IdentificationUralic Language Identification (ULI) 2020 shared task dataset and the Wanca 2017 corpora
This article introduces the Wanca 2017 web corpora from which the sentences written in minor Uralic languages were collected for the test set of the Uralic Language Identification (ULI) 2020 shared task. We describe the …
Language IdentificationLanguages under the influence: Building a database of Uralic languages
Evaluating Transferability of BERT Models on Uralic Languages
Transformer-based language models such as BERT have outperformed previous models on a large number of English benchmarks, but their evaluation is often limited to English or a small number of well-resourced languages. In…
Hyperparameter OptimizationNERPOSGrapheme-Based Cross-Language Forced Alignment: Results with Uralic Languages
Forced alignment is an effective process to speed up linguistic research. However, most forced aligners are language-dependent, and under-resourced languages rarely have enough resources to train an acoustic model for an…