Czech Text Document Corpus v 2.0
This paper introduces "Czech Text Document Corpus v 2.0", a collection of text documents for automatic document classification in Czech language. It is composed of the text documents provided by the Czech News Agency and is freely available for research purposes at http://ctdc.kiv.zcu.cz/. This corpus was created in order to facilitate a straightforward comparison of the document classification approaches on Czech data. It is particularly dedicated to evaluation of multi-label document classification approaches, because one document is usually labelled with more than one label. Besides the information about the document classes, the corpus is also annotated at the morphological layer. This paper further shows the results of selected state-of-the-art methods on this corpus to offer the possibility of an easy comparison with these approaches.
Code (0)
등록된 구현이 없습니다.
Tasks
ClassificationDocument ClassificationGeneral ClassificationSimilar Papers 제목 키워드 기반
Czech Historical Named Entity Corpus v 1.0
As the number of digitized archival documents increases very rapidly, named entity recognition (NER) in historical documents has become very important for information extraction and data mining. For this task an annotate…
named-entity-recognitionNamed Entity RecognitionNamed Entity Recognition (NER)NER+1Deep Neural Networks for Czech Multi-label Document Classification
This paper is focused on automatic multi-label document classification of Czech text documents. The current approaches usually use some pre-processing which can have negative impact (loss of information, additional imple…
ClassificationDocument ClassificationGeneral ClassificationMulti-document multilingual summarization corpus preparation, Part 2: Czech, Hebrew and Spanish
Czech Legal Text Treebank 1.0
We introduce a new member of the family of Prague dependency treebanks. The Czech Legal Text Treebank 1.0 is a morphologically and syntactically annotated corpus of 1,128 sentences. The treebank contains texts from the l…
Transfer Learning for Czech Historical Named Entity Recognition
Nowadays, named entity recognition (NER) achieved excellent results on the standard corpora. However, big issues are emerging with a need for an application in a specific domain, because it requires a suitable annotated …
named-entity-recognitionNamed Entity RecognitionNamed Entity Recognition (NER)NER+2