A Dataset and Strong Baselines for Classification of Czech News Texts
Pre-trained models for Czech Natural Language Processing are often evaluated on purely linguistic tasks (POS tagging, parsing, NER) and relatively simple classification tasks such as sentiment classification or article classification from a single news source. As an alternative, we present CZEch~NEws~Classification~dataset (CZE-NEC), one of the largest Czech classification datasets, composed of news articles from various sources spanning over twenty years, which allows a more rigorous evaluation of such models. We define four classification tasks: news source, news category, inferred author's gender, and day of the week. To verify the task difficulty, we conducted a human evaluation, which revealed that human performance lags behind strong machine-learning baselines built upon pre-trained transformer models. Furthermore, we show that language-specific pre-trained encoder analysis outperforms selected commercially available large-scale generative language models.
Code (1)
Tasks
ArticlesClassificationNERNews ClassificationPOSPOS TaggingSentiment AnalysisSentiment ClassificationSimilar Papers 제목 키워드 기반
SumeCzech: Large Czech News-Based Summarization Dataset
Czech Text Document Corpus v 2.0
This paper introduces "Czech Text Document Corpus v 2.0", a collection of text documents for automatic document classification in Czech language. It is composed of the text documents provided by the Czech News Agency and…
ClassificationDocument ClassificationGeneral ClassificationBenchmark Dataset for Propaganda Detection in Czech Newspaper Texts
Propaganda of various pressure groups ranging from big economies to ideological blocks is often presented in a form of objective newspaper texts. However, the real objectivity is here shaded with the support of imbalance…
ArticlesPropaganda detectionCUNI Systems in WMT21: Revisiting Backtranslation Techniques for English-Czech NMT
We describe our two NMT systems submitted to the WMT2021 shared task in English-Czech news translation: CUNI-DocTransformer (document-level CUBBITT) and CUNI-Marian-Baselines. We improve the former with a better sentence…
NMTSegmentationSentenceSentence segmentation+1Text Summarization of Czech News Articles Using Named Entities
The foundation for the research of summarization in the Czech language was laid by the work of Straka et al. (2018). They published the SumeCzech, a large Czech news-based summarization dataset, and proposed several base…
Abstractive Text SummarizationArticlesExtractive SummarizationSentence+1