paper-with-me

홈 › Papers

A data-based classification of Slavic languages: Indices of qualitative variation applied to grapheme frequencies

2015-04-14 · Michaela Koscová, Ján Macutek, Emmerich Kelih

The Ord's graph is a simple graphical method for displaying frequency distributions of data or theoretical distributions in the two-dimensional plane. Its coordinates are proportions of the first three moments, either empirical or theoretical ones. A modification of the Ord's graph based on proportions of indices of qualitative variation is presented. Such a modification makes the graph applicable also to data of categorical character. In addition, the indices are normalized with values between 0 and 1, which enables comparing data files divided into different numbers of categories. Both the original and the new graph are used to display grapheme frequencies in eleven Slavic languages. As the original Ord's graph requires an assignment of numbers to the categories, graphemes were ordered decreasingly according to their frequencies. Data were taken from parallel corpora, i.e., we work with grapheme frequencies from a Russian novel and its translations to ten other Slavic languages. Then, cluster analysis is applied to the graph coordinates. While the original graph yields results which are not linguistically interpretable, the modification reveals meaningful relations among the languages.

📄 PDF Abstract BibTeX arXiv:1504.03608

Code (0)

등록된 구현이 없습니다.

Tasks

General Classification

Similar Papers 제목 키워드 기반

The Second Cross-Lingual Challenge on Recognition, Normalization, Classification, and Linking of Named Entities across Slavic Languages

2019-08-01 · WS 2019 8 · Jakub Piskorski, Laska Laskova, Micha{\l} Marci{\'n}czuk, Lidia Pivovarova 외

We describe the Second Multilingual Named Entity Challenge in Slavic languages. The task is recognizing mentions of named entities in Web documents, their normalization, and cross-lingual linking. The Challenge was organ…

Cross-Lingual Entity LinkingEntity Linkingnamed-entity-recognitionNamed Entity Recognition+1

Slav-NER: the 3rd Cross-lingual Challenge on Recognition, Normalization, Classification, and Linking of Named Entities across Slavic Languages

2021-04-01 · EACL (BSNLP) 2021 4 · Jakub Piskorski, Bogdan Babych, Zara Kancheva, Olga Kanishcheva 외

This paper describes Slav-NER: the 3rd Multilingual Named Entity Challenge in Slavic languages. The tasks involve recognizing mentions of named entities in Web documents, normalization of the names, and cross-lingual lin…

Cross-Lingual Entity LinkingEntity Linkingnamed-entity-recognitionNamed Entity Recognition+2

State of the Art in Text Classification for South Slavic Languages: Fine-Tuning or Prompting?

2025-11-11 · Taja Kuzman Pungeršek, Peter Rupnik, Ivan Porupski, Vuk Dinić 외 arxiv

Until recently, fine-tuned BERT-like models provided state-of-the-art performance on text classification tasks. With the rise of instruction-tuned decoder-only models, commonly known as large language models (LLMs), the …

Text Classification

Universal Dependencies for Serbian in Comparison with Croatian and Other Slavic Languages

2017-04-01 · WS 2017 4 · Tanja Samard{\v{z}}i{\'c}, Mirjana Starovi{\'c}, {\v{Z}}eljko Agi{\'c}, Nikola Ljube{\v{s}}i{\'c}

The paper documents the procedure of building a new Universal Dependencies (UDv2) treebank for Serbian starting from an existing Croatian UDv1 treebank and taking into account the other Slavic UD annotation guidelines. W…

Tuning Multilingual Transformers for Named Entity Recognition on Slavic Languages

2019-01-30 · Conference: Proceedings of the 7th Workshop on Balto-Slavic Natural Language Processing 2019 1 · Mikhail Arkhipov, Maria Trofimova, Yuri Kuratov, Alexey Sorokin

Our paper addresses the problem of multilingual named entity recognition on the material of 4 languages: Russian, Bulgarian, Czech and Polish. We solve this task using the BERT model. We use a hundred languages multiling…

Multilingual Named Entity Recognitionnamed-entity-recognitionNamed Entity RecognitionNamed Entity Recognition (NER)+2