paper-with-me

Papers

Russian Web Tables: A Public Corpus of Web Tables for Russian Language Based on Wikipedia

2022-10-03 · Platon Fedorov, Alexey Mironov, George Chernishev

Corpora that contain tabular data such as WebTables are a vital resource for the academic community. Essentially, they are the backbone of any modern research in information management. They are used for various tasks of data extraction, knowledge base construction, question answering, column semantic type detection and many other. Such corpora are useful not only as a source of data, but also as a base for building test datasets. So far, there were no such corpora for the Russian language and this seriously hindered research in the aforementioned areas. In this paper, we present the first corpus of Web tables created specifically out of Russian language material. It was built via a special toolkit we have developed to crawl the Russian Wikipedia. Both the corpus and the toolkit are open-source and publicly available. Finally, we present a short study that describes Russian Wikipedia tables and their statistics.

📄 PDF Abstract BibTeX arXiv:2210.06353

Code (3)

https://gitlab.com/unidata-labs/ru-wiki-tables-backend 공식 구현
https://gitlab.com/unidata-labs/ru-wiki-tables-dataset 공식 구현
https://gitlab.com/unidata-labs/ru-wiki-tables-frontend 공식 구현

Tasks

Knowledge Base ConstructionManagementQuestion Answering

Methods 이 논문이 사용한 방법론

Test 설명 없음
BASE 설명 없음

Similar Papers 제목 키워드 기반

Analysis of the quotation corpus of the Russian Wiktionary

2020-01-20 · A. Smirnov, T. Levashova, A. Karpov, I. Kipyatkova 외

The quantitative evaluation of quotations in the Russian Wiktionary was performed using the developed Wiktionary parser. It was found that the number of quotations in the dictionary is growing fast (51.5 thousands in 201…

Russian-Language Multimodal Dataset for Automatic Summarization of Scientific Papers

2024-05-13 · Alena Tsanda, Elena Bruches

The paper discusses the creation of a multimodal dataset of Russian-language scientific papers and testing of existing language models for the task of automatic text summarization. A feature of the dataset is its multimo…

Text Summarization

RuCoCo: a new Russian corpus with coreference annotation

2022-06-10 · Vladimir Dobrovolskii, Mariia Michurina, Alexandra Ivoylova

We present a new corpus with coreference annotation, Russian Coreference Corpus (RuCoCo). The goal of RuCoCo is to obtain a large number of annotated texts while maintaining high inter-annotator agreement. RuCoCo contain…

The Algorithmic Inflection of Russian and Generation of Grammatically Correct Text

2017-06-08 · T. M. Sadykov, T. A. Zhukov

We present a deterministic algorithm for Russian inflection. This algorithm is implemented in a publicly available web-service www.passare.ru which provides functions for inflection of single words, word matching and syn…

The Algorithmic Inflection and Morphological Variability of Russian

2021-11-16 · ACL ARR November 2021 11 · Anonymous

We present a set of deterministic algorithms for Russian inflection and automated text synthesis. These algorithms are implemented in a publicly available web-service www.passare.ru. This service provides functions for i…