Bootstrapping Multilingual Metadata Extraction: A Showcase in Cyrillic
Applications based on scholarly data are of ever increasing importance. This results in disadvantages for areas where high-quality data and compatible systems are not available, such as non-English publications. To advance the mitigation of this imbalance, we use Cyrillic script publications from the CORE collection to create a high-quality data set for metadata extraction. We utilize our data for training and evaluating sequence labeling models to extract title and author information. Retraining GROBID on our data, we observe significant improvements in terms of precision and recall and achieve even better results with a self developed model. We make our data set covering over 15,000 publications as well as our source code freely available.
Code (1)
Similar Papers 제목 키워드 기반
PulseBench-Tab: A Multilingual Benchmark for Table Extraction with Graph-Based Evaluation
We introduce PulseBench-Tab, an open multilingual benchmark for evaluating table extraction from document images. The benchmark comprises 1,820 human-annotated tables spanning 9 languages and 4 scripts (Latin, CJK, Arabi…
Adapting the LodView RDF Browser for Navigation over the Multilingual Linguistic Linked Open Data Cloud
The paper is dedicated to the use of LodView for navigation over the multilingual Linguistic Linked Open Data cloud. First, we define the class of Pubby-like tools, that LodView belongs to, and clarify the relation of th…
MathCyrillic-MNIST: a Cyrillic Version of the MNIST Dataset
This paper presents a new handwritten dataset, Cyrillic-MNIST, a Cyrillic version of the MNIST dataset, comprising of 121,234 samples of 42 Cyrillic letters. The performance of Cyrillic-MNIST is evaluated using standard …
Classification of Handwritten Names of Cities and Handwritten Text Recognition using Various Deep Learning Models
This article discusses the problem of handwriting recognition in Kazakh and Russian languages. This area is poorly studied since in the literature there are almost no works in this direction. We have tried to describe va…
Handwriting RecognitionHandwritten Text RecognitionUzbek Cyrillic-Latin-Cyrillic Machine Transliteration
In this paper, we introduce a data-driven approach to transliterating Uzbek dictionary words from the Cyrillic script into the Latin script, and vice versa. We heuristically align characters of words in the source script…
Transliteration