paper-with-me

홈 › Papers

skLEP: A Slovak General Language Understanding Benchmark

2025-06-26 · Marek Šuppa, Andrej Ridzik, Daniel Hládek, Tomáš Javůrek, Viktória Ondrejová, Kristína Sásiková, Martin Tamajka, Marián Šimko

In this work, we introduce skLEP, the first comprehensive benchmark specifically designed for evaluating Slovak natural language understanding (NLU) models. We have compiled skLEP to encompass nine diverse tasks that span token-level, sentence-pair, and document-level challenges, thereby offering a thorough assessment of model capabilities. To create this benchmark, we curated new, original datasets tailored for Slovak and meticulously translated established English NLU resources. Within this paper, we also present the first systematic and extensive evaluation of a wide array of Slovak-specific, multilingual, and English pre-trained language models using the skLEP tasks. Finally, we also release the complete benchmark data, an open-source toolkit facilitating both fine-tuning and evaluation of models, and a public leaderboard at https://github.com/slovak-nlp/sklep in the hopes of fostering reproducibility and drive future research in Slovak NLU.

📄 PDF Abstract BibTeX arXiv:2506.21508

Code (0)

등록된 구현이 없습니다.

Tasks

Natural Language UnderstandingSentence

Similar Papers 제목 키워드 기반

SlovakBERT: Slovak Masked Language Model

2021-09-30 · Matúš Pikuliak, Štefan Grivalský, Martin Konôpka, Miroslav Blšták 외

We introduce a new Slovak masked language model called SlovakBERT. This is to our best knowledge the first paper discussing Slovak transformers-based language models. We evaluate our model on several NLP tasks and achiev…

Language ModelingLanguage ModellingmodelPart-Of-Speech Tagging+2

SkMTEB: Slovak Massive Text Embedding Benchmark and Model Adaptation

2026-06-11 · Marek Šuppa, Andrej Ridzik, Daniel Hládek, Natália Kňažeková 외 arxiv

We introduce SkMTEB, the first comprehensive MTEB-style text embedding benchmark for Slovak, a low-resource West Slavic language, comprising 31 datasets across 7 task types -- nearly 4$\times$ the depth of existing multi…

WikiGoldSK: Annotated Dataset, Baselines and Few-Shot Learning Experiments for Slovak Named Entity Recognition

2023-04-08 · Dávid Šuba, Marek Šuppa, Jozef Kubík, Endre Hamerlik 외

Named Entity Recognition (NER) is a fundamental NLP tasks with a wide range of practical applications. The performance of state-of-the-art NER methods depends on high quality manually anotated datasets which still do not…

Few-Shot Learningnamed-entity-recognitionNamed Entity RecognitionNamed Entity Recognition (NER)+1

SlovKE: A Large-Scale Dataset and LLM Evaluation for Slovak Keyphrase Extraction

2026-03-16 · David Števaňák, Marek Šuppa arxiv

Keyphrase extraction for morphologically rich, low-resource languages remains understudied, largely due to the scarcity of suitable evaluation datasets. We address this gap for Slovak by constructing a dataset of 227,432…

Keyphrase Extraction

TUKE-BNews-SK: Slovak Broadcast News Corpus Construction and Evaluation

2014-05-01 · LREC 2014 5 · Mat{\'u}{\v{s}} Pleva, Jozef Juh{\'a}r

This article presents an overview of the existing acoustical corpuses suitable for broadcast news automatic transcription task in the Slovak language. The TUKE-BNews-SK database created in our department was built to sup…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)speech-recognitionSpeech Recognition