paper-with-me

Papers

Evaluating Transferability of BERT Models on Uralic Languages

2021-09-13 · ACL (IWCLUL) 2021 9 · Judit Ács, Dániel Lévai, András Kornai

Transformer-based language models such as BERT have outperformed previous models on a large number of English benchmarks, but their evaluation is often limited to English or a small number of well-resourced languages. In this work, we evaluate monolingual, multilingual, and randomly initialized language models from the BERT family on a variety of Uralic languages including Estonian, Finnish, Hungarian, Erzya, Moksha, Karelian, Livvi, Komi Permyak, Komi Zyrian, Northern S\'ami, and Skolt S\'ami. When monolingual models are available (currently only et, fi, hu), these perform better on their native language, but in general they transfer worse than multilingual models or models of genetically unrelated languages that share the same character set. Remarkably, straightforward transfer of high-resource models, even without special efforts toward hyperparameter optimization, yields what appear to be state of the art POS and NER tools for the minority Uralic languages where there is sufficient data for finetuning.

📄 PDF Abstract BibTeX arXiv:2109.06327

Code (1)

juditacs/uralic_eval 공식 구현 pytorch

Tasks

Hyperparameter OptimizationNERPOS

Methods 이 논문이 사용한 방법론

Multi-Head Attention 설명 없음
Attention 설명 없음
Linear Layer A Linear Layer is a projection $\mathbf{XW + b}$.
Linear Warmup With Linear Decay Linear Warmup With Linear Decay is a learning rate schedule in which we increase the learning rate linearly for $n$ updates and then linearly decay afterwards.
Weight Decay 설명 없음
Attention Dropout Attention Dropout is a type of dropout used in attention-based architectures, where elements are randomly dropped out of the…
Dropout Dropout is a regularization technique for neural networks that drops a unit (along with connections) at training time with a specified probability $p$ (a common value is…
Layer Normalization Unlike batch normalization, Layer Normalization directly estimates the normalization statistics from the summed inputs…

Similar Papers 제목 키워드 기반

UralicNLP: An NLP Library for Uralic Languages

2019-05-09 · Journal of Open Source Software 2019 5 · Mika Hämäläinen

UralicNLP is a natural language processing library for small Uralic languages. It can produce morphological analysis, generate morphological forms, lemmatize words and give lexical information about words in Uralic langu…

Morphological Analysis

Evaluating OpenAI GPT Models for Translation of Endangered Uralic Languages: A Comparison of Reasoning and Non-Reasoning Architectures

2025-12-18 · Yehor Tereshchenko, Mika Hämäläinen, Svitlana Myroniuk arxiv

The evaluation of Large Language Models (LLMs) for translation tasks has primarily focused on high-resource languages, leaving a significant gap in understanding their performance on low-resource and endangered languages…

Uralic Language Identification (ULI) 2020 shared task dataset and the Wanca 2017 corpus

2020-08-27 · Tommi Jauhiainen, Heidi Jauhiainen, Niko Partanen, Krister Lindén

This article introduces the Wanca 2017 corpus of texts crawled from the internet from which the sentences in rare Uralic languages for the use of the Uralic Language Identification (ULI) 2020 shared task were collected. …

Language Identification

Uralic Language Identification (ULI) 2020 shared task dataset and the Wanca 2017 corpora

2020-12-01 · VarDial (COLING) 2020 12 · Tommi Jauhiainen, Heidi Jauhiainen, Niko Partanen, Krister Lindén

This article introduces the Wanca 2017 web corpora from which the sentences written in minor Uralic languages were collected for the test set of the Uralic Language Identification (ULI) 2020 shared task. We describe the …

Language Identification

Languages under the influence: Building a database of Uralic languages

2017-01-01 · WS 2017 1 · Eszter Simon, Nikolett Mus