paper-with-me

Papers

Analysis of the quotation corpus of the Russian Wiktionary

2020-01-20 · A. Smirnov, T. Levashova, A. Karpov, I. Kipyatkova, A. Ronzhin, A. Krizhanovsky, N. Krizhanovsky

The quantitative evaluation of quotations in the Russian Wiktionary was performed using the developed Wiktionary parser. It was found that the number of quotations in the dictionary is growing fast (51.5 thousands in 2011, 62 thousands in 2012). These quotations were extracted and saved in the relational database of a machine-readable dictionary. For this database, tables related to the quotations were designed. A histogram of distribution of quotations of literary works written in different years was built. It was made an attempt to explain the characteristics of the histogram by associating it with the years of the most popular and cited (in the Russian Wiktionary) writers of the nineteenth century. It was found that more than one-third of all the quotations (the example sentences) contained in the Russian Wiktionary are taken by the editors of a Wiktionary entry from the Russian National Corpus.

📄 PDF Abstract BibTeX arXiv:2002.00734

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

The comparison of Wiktionary thesauri transformed into the machine-readable format

2010-06-25 · A. A. Krizhanovsky

Wiktionary is a unique, peculiar, valuable and original resource for natural language processing (NLP). The paper describes an open-source Wiktionary parser: its architecture and requirements followed by a description of…

Related terms search based on WordNet / Wiktionary and its application in Ontology Matching

2009-07-13 · A. A. Krizhanovsky, Feiyu Lin

A set of ontology matching algorithms (for finding correspondences between concepts) is based on a thesaurus that provides the source data for the semantic distance calculations. In this wiki era, new resources may sprin…

Ontology Matching

The Project Dialogism Novel Corpus: A Dataset for Quotation Attribution in Literary Texts

2022-04-12 · LREC 2022 6 · Krishnapriya Vishnubhotla, Adam Hammond, Graeme Hirst

We present the Project Dialogism Novel Corpus, or PDNC, an annotated dataset of quotations for English literary texts. PDNC contains annotations for 35,978 quotations across 22 full-length novels, and is by an order of m…

Referring Expression

Quotations, Coreference Resolution, and Sentiment Annotations in Croatian News Articles: An Exploratory Study

2022-12-14 · Jelena Sarajlić, Gaurish Thakkar, Diego Alves, Nives Mikelic Preradović

This paper presents a corpus annotated for the task of direct-speech extraction in Croatian. The paper focuses on the annotation of the quotation, co-reference resolution, and sentiment annotation in SETimes news corpus …

Articlescoreference-resolutionCoreference ResolutionSpeech Extraction

Semi-automatic methods for adding words to the dictionary of VepKar corpus based on inflectional rules extracted from Wiktionary

2020-01-14 · Natalia Krizhanovsky, Andrew Krizhanovsky

The article describes a technique for using English Wiktionary inflection tables for generating word forms for Veps verbs and nominals in the Open corpus of Veps and Karelian languages. The information concerning Karelia…