paper-with-me

Papers

UralicNLP: An NLP Library for Uralic Languages

2019-05-09 · Journal of Open Source Software 2019 5 · Mika Hämäläinen

UralicNLP is a natural language processing library for small Uralic languages. It can produce morphological analysis, generate morphological forms, lemmatize words and give lexical information about words in Uralic languages. At the time of writing, the following languages are supported: Skolt Sami, Ingrian, Meadow & Eastern Mari, Votic, Olonets-Karelian, Erzya, Moksha, Hill Mari, Udmurt, Tundra Nenets, Komi-Permyak and Finnish. This information originates from FST tools and dictionaries developed in the Giellatekno infrastructure. Currently, UralicNLP uses the nightly builds for languages supported by Apertium and less frequently updated FSTs and CGs for the other languages.

📄 PDF Abstract BibTeX

Code (1)

mikahama/uralicNLP 공식 구현

Tasks

Morphological Analysis

Similar Papers 제목 키워드 기반

Uralic Language Identification (ULI) 2020 shared task dataset and the Wanca 2017 corpus

2020-08-27 · Tommi Jauhiainen, Heidi Jauhiainen, Niko Partanen, Krister Lindén

This article introduces the Wanca 2017 corpus of texts crawled from the internet from which the sentences in rare Uralic languages for the use of the Uralic Language Identification (ULI) 2020 shared task were collected. …

Language Identification

Uralic Language Identification (ULI) 2020 shared task dataset and the Wanca 2017 corpora

2020-12-01 · VarDial (COLING) 2020 12 · Tommi Jauhiainen, Heidi Jauhiainen, Niko Partanen, Krister Lindén

This article introduces the Wanca 2017 web corpora from which the sentences written in minor Uralic languages were collected for the test set of the Uralic Language Identification (ULI) 2020 shared task. We describe the …

Language Identification

Languages under the influence: Building a database of Uralic languages

2017-01-01 · WS 2017 1 · Eszter Simon, Nikolett Mus

Evaluating Transferability of BERT Models on Uralic Languages

2021-09-13 · ACL (IWCLUL) 2021 9 · Judit Ács, Dániel Lévai, András Kornai

Transformer-based language models such as BERT have outperformed previous models on a large number of English benchmarks, but their evaluation is often limited to English or a small number of well-resourced languages. In…

Hyperparameter OptimizationNERPOS

Grapheme-Based Cross-Language Forced Alignment: Results with Uralic Languages

2021-05-01 · NoDaLiDa 2021 5 · Juho Leinonen, Sami Virpioja, Mikko Kurimo

Forced alignment is an effective process to speed up linguistic research. However, most forced aligners are language-dependent, and under-resourced languages rarely have enough resources to train an acoustic model for an…