paper-with-me

홈 › Papers

Deep Neural Networks for Czech Multi-label Document Classification

2017-01-13 · Ladislav Lenc, Pavel Král

This paper is focused on automatic multi-label document classification of Czech text documents. The current approaches usually use some pre-processing which can have negative impact (loss of information, additional implementation work, etc). Therefore, we would like to omit it and use deep neural networks that learn from simple features. This choice was motivated by their successful usage in many other machine learning fields. Two different networks are compared: the first one is a standard multi-layer perceptron, while the second one is a popular convolutional network. The experiments on a Czech newspaper corpus show that both networks significantly outperform baseline method which uses a rich set of features with maximum entropy classifier. We have also shown that convolutional network gives the best results.

📄 PDF Abstract BibTeX arXiv:1701.03849

Code (0)

등록된 구현이 없습니다.

Tasks

ClassificationDocument ClassificationGeneral Classification

Similar Papers 제목 키워드 기반

Czech Text Document Corpus v 2.0

2017-10-06 · LREC 2018 5 · Pavel Král, Ladislav Lenc

This paper introduces "Czech Text Document Corpus v 2.0", a collection of text documents for automatic document classification in Czech language. It is composed of the text documents provided by the Czech News Agency and…

ClassificationDocument ClassificationGeneral Classification

Word Embeddings for Multi-label Document Classification

2017-09-01 · RANLP 2017 9 · Ladislav Lenc, Pavel Kr{\'a}l

In this paper, we analyze and evaluate word embeddings for representation of longer texts in the multi-label classification scenario. The embeddings are used in three convolutional neural network topologies. The experime…

ClassificationDocument ClassificationGeneral ClassificationMulti-Label Classification+4

CzechDocs: A Multiway Parallel Dataset of Formatted Documents for Minority Languages in Czechia

2026-06-18 · Josef Jon, Ondřej Bojar arxiv

We present CzechDocs, a multiway parallel dataset of formatted documents (HTML, DOCX, and PDF) covering Czech and minority languages used in Czechia-primarily Ukrainian and English, with smaller portions of Vietnamese, R…

Machine Translation

CWRCzech: 100M Query-Document Czech Click Dataset and Its Application to Web Relevance Ranking

2024-05-31 · Josef Vonášek, Milan Straka, Rostislav Krč, Lenka Lasoňová 외

We present CWRCzech, Click Web Ranking dataset for Czech, a 100M query-document Czech click dataset for relevance ranking with user behavior data collected from search engine logs of Seznam$.$cz. To the best of our knowl…

Czech Historical Named Entity Corpus v 1.0

2020-05-01 · LREC 2020 5 · Helena Hubkov{\'a}, Pavel Kral, Eva Pettersson

As the number of digitized archival documents increases very rapidly, named entity recognition (NER) in historical documents has become very important for information extraction and data mining. For this task an annotate…

named-entity-recognitionNamed Entity RecognitionNamed Entity Recognition (NER)NER+1