paper-with-me

Papers

Pretraining and Benchmarking Modern Encoders for Latvian

2026-03-16 · Arturs Znotins arxiv

Encoder-only transformers remain essential for practical NLP tasks. While recent advances in multilingual models have improved cross-lingual capabilities, low-resource languages such as Latvian remain underrepresented in pretraining corpora, and few monolingual Latvian encoders currently exist. We address this gap by pretraining a suite of Latvian-specific encoders based on RoBERTa, DeBERTaV3, and ModernBERT architectures, including long-context variants, and evaluating them across a diverse set of Latvian diagnostic and linguistic benchmarks. Our models are competitive with existing monolingual and multilingual encoders while benefiting from recent architectural and efficiency advances. Our best model, lv-deberta-base (111M parameters), achieves the strongest overall performance, outperforming larger multilingual baselines and prior Latvian-specific encoders. We release all pretrained models and evaluation resources to support further research and practical applications in Latvian NLP.

📄 PDF Abstract BibTeX arXiv:2603.15005

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Pretraining and Fine-Tuning Strategies for Sentiment Analysis of Latvian Tweets

2020-10-23 · Gaurish Thakkar, Marcis Pinnis

In this paper, we present various pre-training strategies that aid in im-proving the accuracy of the sentiment classification task. We, at first, pre-trainlanguage representation models using these strategies and then fi…

Sentiment AnalysisSentiment Classification

LAG-MMLU: Benchmarking Frontier LLM Understanding in Latvian and Giriama

2025-03-14 · Naome A. Etori, Kevin Lu, Randu Karisa, Arturs Kanepajs

As large language models (LLMs) rapidly advance, evaluating their performance is critical. LLMs are trained on multilingual data, but their reasoning abilities are mainly evaluated using English datasets. Hence, robust e…

BenchmarkingMMLU

Latvian National Corpora Collection – Korpuss.lv

2022-06-01 · LREC 2022 6 · Baiba Saulite, Roberts Darģis, Normunds Gruzitis, Ilze Auzina 외

LNCC is a diverse collection of Latvian language corpora representing both written and spoken language and is useful for both linguistic research and language modelling. The collection is intended to cover diverse Latvia…

Cultural Vocal Bursts Intensity PredictionLanguage Modelling

Masked Capsule Autoencoders

2024-03-07 · Miles Everett, Mingjun Zhong, Georgios Leontidis

We propose Masked Capsule Autoencoders (MCAE), the first Capsule Network that utilises pretraining in a modern self-supervised paradigm, specifically the masked image modelling framework. Capsule Networks have emerged as…

Decoder

LaVA – Latvian Language Learner corpus

2022-06-01 · LREC 2022 6 · Roberts Darģis, Ilze Auziņa, Inga Kaija, Kristīne Levāne-Petrova 외

This paper presents the Latvian Language Learner Corpus (LaVA) developed at the Institute of Mathematics and Computer Science, University of Latvia. LaVA corpus contains 1015 essays (190k tokens and 790k characters exclu…