paper-with-me

홈 › Papers

Bootstrapping Multilingual Metadata Extraction: A Showcase in Cyrillic

2021-06-01 · NAACL (sdp) 2021 6 · Johan Krause, Igor Shapiro, Tarek Saier, Michael Färber

Applications based on scholarly data are of ever increasing importance. This results in disadvantages for areas where high-quality data and compatible systems are not available, such as non-English publications. To advance the mitigation of this imbalance, we use Cyrillic script publications from the CORE collection to create a high-quality data set for metadata extraction. We utilize our data for training and evaluating sequence labeling models to extract title and author information. Retraining GROBID on our data, we observe significant improvements in terms of precision and recall and achieve even better results with a self developed model. We make our data set covering over 15,000 publications as well as our source code freely available.

📄 PDF Abstract BibTeX

Code (1)

illdepence/sdp2021 공식 구현 tf

Similar Papers 제목 키워드 기반

PulseBench-Tab: A Multilingual Benchmark for Table Extraction with Graph-Based Evaluation

2026-04-21 · Ritvik Pandey, Sid Manchkanti, Mohammed Wazir Adain, Mohammed Hadi 외 arxiv

We introduce PulseBench-Tab, an open multilingual benchmark for evaluating table extraction from document images. The benchmark comprises 1,820 human-annotated tables spanning 9 languages and 4 scripts (Latin, CJK, Arabi…

Adapting the LodView RDF Browser for Navigation over the Multilingual Linguistic Linked Open Data Cloud

2022-08-28 · Alexander Kirillovich, Konstantin Nikolaev

The paper is dedicated to the use of LodView for navigation over the multilingual Linguistic Linked Open Data cloud. First, we define the class of Pubby-like tools, that LodView belongs to, and clarify the relation of th…

Math

Cyrillic-MNIST: a Cyrillic Version of the MNIST Dataset

2022-06-01 · LREC 2022 6 · Bolat Tleubayev, Zhanel Zhexenova, Kenessary Koishybay, Anara Sandygulova

This paper presents a new handwritten dataset, Cyrillic-MNIST, a Cyrillic version of the MNIST dataset, comprising of 121,234 samples of 42 Cyrillic letters. The performance of Cyrillic-MNIST is evaluated using standard …

Classification of Handwritten Names of Cities and Handwritten Text Recognition using Various Deep Learning Models

2021-02-09 · Daniyar Nurseitov, Kairat Bostanbekov, Maksat Kanatov, Anel Alimova 외

This article discusses the problem of handwriting recognition in Kazakh and Russian languages. This area is poorly studied since in the literature there are almost no works in this direction. We have tried to describe va…

Handwriting RecognitionHandwritten Text Recognition

Uzbek Cyrillic-Latin-Cyrillic Machine Transliteration

2021-01-13 · B. Mansurov, A. Mansurov

In this paper, we introduce a data-driven approach to transliterating Uzbek dictionary words from the Cyrillic script into the Latin script, and vice versa. We heuristically align characters of words in the source script…

Transliteration