paper-with-me

홈 › Papers

Czech Text Document Corpus v 2.0

2017-10-06 · LREC 2018 5 · Pavel Král, Ladislav Lenc

This paper introduces "Czech Text Document Corpus v 2.0", a collection of text documents for automatic document classification in Czech language. It is composed of the text documents provided by the Czech News Agency and is freely available for research purposes at http://ctdc.kiv.zcu.cz/. This corpus was created in order to facilitate a straightforward comparison of the document classification approaches on Czech data. It is particularly dedicated to evaluation of multi-label document classification approaches, because one document is usually labelled with more than one label. Besides the information about the document classes, the corpus is also annotated at the morphological layer. This paper further shows the results of selected state-of-the-art methods on this corpus to offer the possibility of an easy comparison with these approaches.

📄 PDF Abstract BibTeX arXiv:1710.02365

Code (0)

등록된 구현이 없습니다.

Tasks

ClassificationDocument ClassificationGeneral Classification

Similar Papers 제목 키워드 기반

Czech Historical Named Entity Corpus v 1.0

2020-05-01 · LREC 2020 5 · Helena Hubkov{\'a}, Pavel Kral, Eva Pettersson

As the number of digitized archival documents increases very rapidly, named entity recognition (NER) in historical documents has become very important for information extraction and data mining. For this task an annotate…

named-entity-recognitionNamed Entity RecognitionNamed Entity Recognition (NER)NER+1

Deep Neural Networks for Czech Multi-label Document Classification

2017-01-13 · Ladislav Lenc, Pavel Král

This paper is focused on automatic multi-label document classification of Czech text documents. The current approaches usually use some pre-processing which can have negative impact (loss of information, additional imple…

ClassificationDocument ClassificationGeneral Classification

Multi-document multilingual summarization corpus preparation, Part 2: Czech, Hebrew and Spanish

2013-08-01 · WS 2013 8 · Michael Elhadad, Mir, Sabino a-Jim{\'e}nez, Josef Steinberger 외
Document SummarizationMulti-Document Summarization

Czech Legal Text Treebank 1.0

2016-05-01 · LREC 2016 5 · Vincent Kr{\'\i}{\v{z}}, Barbora Hladk{\'a}, Zde{\v{n}}ka Ure{\v{s}}ov{\'a}

We introduce a new member of the family of Prague dependency treebanks. The Czech Legal Text Treebank 1.0 is a morphologically and syntactically annotated corpus of 1,128 sentences. The treebank contains texts from the l…

Transfer Learning for Czech Historical Named Entity Recognition

2021-09-01 · RANLP 2021 9 · Helena Hubková, Pavel Kral

Nowadays, named entity recognition (NER) achieved excellent results on the standard corpora. However, big issues are emerging with a need for an application in a specific domain, because it requires a suitable annotated …

named-entity-recognitionNamed Entity RecognitionNamed Entity Recognition (NER)NER+2