paper-with-me

Papers

BERTić -- The Transformer Language Model for Bosnian, Croatian, Montenegrin and Serbian

2021-04-19 · Nikola Ljubešić, Davor Lauc

In this paper we describe a transformer model pre-trained on 8 billion tokens of crawled text from the Croatian, Bosnian, Serbian and Montenegrin web domains. We evaluate the transformer model on the tasks of part-of-speech tagging, named-entity-recognition, geo-location prediction and commonsense causal reasoning, showing improvements on all tasks over state-of-the-art models. For commonsense reasoning evaluation, we introduce COPA-HR -- a translation of the Choice of Plausible Alternatives (COPA) dataset into Croatian. The BERTi\'c model is made available for free usage and further task-specific fine-tuning through HuggingFace.

📄 PDF Abstract BibTeX arXiv:2104.09243

Code (0)

등록된 구현이 없습니다.

Tasks

Commonsense Causal ReasoningLanguage ModelingLanguage Modellingnamed-entity-recognitionNamed Entity RecognitionNamed Entity Recognition (NER)Part-Of-Speech TaggingTranslation

Similar Papers 제목 키워드 기반

BERTić - The Transformer Language Model for Bosnian, Croatian, Montenegrin and Serbian

2021-04-01 · EACL (BSNLP) 2021 4 · Nikola Ljubešić, Davor Lauc

In this paper we describe a transformer model pre-trained on 8 billion tokens of crawled text from the Croatian, Bosnian, Serbian and Montenegrin web domains. We evaluate the transformer model on the tasks of part-of-spe…

Commonsense Causal ReasoningLanguage ModelingLanguage Modellingnamed-entity-recognition+4

Language Models on a Diet: Cost-Efficient Development of Encoders for Closely-Related Languages via Additional Pretraining

2024-04-08 · Nikola Ljubešić, Vít Suchomel, Peter Rupnik, Taja Kuzman 외

The world of language models is going through turbulent times, better and ever larger models are coming out at an unprecedented speed. However, we argue that, especially for the scientific community, encoder models of up…

CLASSLA-web: Comparable Web Corpora of South Slavic Languages Enriched with Linguistic and Genre Annotation

2024-03-19 · Nikola Ljubešić, Taja Kuzman

This paper presents a collection of highly comparable web corpora of Slovenian, Croatian, Bosnian, Montenegrin, Serbian, Macedonian, and Bulgarian, covering thereby the whole spectrum of official languages in the South S…

Articles

The Growing Gains and Pains of Iterative Web Corpora Crawling: Insights from South Slavic CLASSLA-web 2.0 Corpora

2026-01-16 · Taja Kuzman Pungeršek, Peter Rupnik, Vít Suchomel, Nikola Ljubešić arxiv

Crawling national top-level domains has proven to be highly effective for collecting texts in less-resourced languages. This approach has been recently used for South Slavic languages and resulted in the largest general …

bs,hr,srWaC - Web Corpora of Bosnian, Croatian and Serbian

2014-04-01 · WS 2014 4 · Nikola Ljube{\v{s}}i{\'c}, Filip Klubi{\v{c}}ka
Language IdentificationLanguage Modelling