paper-with-me

홈 › Papers

A Dataset and Strong Baselines for Classification of Czech News Texts

2023-07-20 · Hynek Kydlíček, Jindřich Libovický

Pre-trained models for Czech Natural Language Processing are often evaluated on purely linguistic tasks (POS tagging, parsing, NER) and relatively simple classification tasks such as sentiment classification or article classification from a single news source. As an alternative, we present CZEch~NEws~Classification~dataset (CZE-NEC), one of the largest Czech classification datasets, composed of news articles from various sources spanning over twenty years, which allows a more rigorous evaluation of such models. We define four classification tasks: news source, news category, inferred author's gender, and day of the week. To verify the task difficulty, we conducted a human evaluation, which revealed that human performance lags behind strong machine-learning baselines built upon pre-trained transformer models. Furthermore, we show that language-specific pre-trained encoder analysis outperforms selected commercially available large-scale generative language models.

📄 PDF Abstract BibTeX arXiv:2307.10666

Code (1)

hynky1999/czech-news-classification-dataset 공식 구현

Tasks

ArticlesClassificationNERNews ClassificationPOSPOS TaggingSentiment AnalysisSentiment Classification

Similar Papers 제목 키워드 기반

SumeCzech: Large Czech News-Based Summarization Dataset

2018-05-01 · LREC 2018 5 · Milan Straka, Nikita Mediankin, Tom Kocmi, Zden{\v{e}}k {\v{Z}}abokrtsk{\'y} 외
Document SummarizationMachine TranslationSentence Compression

Czech Text Document Corpus v 2.0

2017-10-06 · LREC 2018 5 · Pavel Král, Ladislav Lenc

This paper introduces "Czech Text Document Corpus v 2.0", a collection of text documents for automatic document classification in Czech language. It is composed of the text documents provided by the Czech News Agency and…

ClassificationDocument ClassificationGeneral Classification

Benchmark Dataset for Propaganda Detection in Czech Newspaper Texts

2019-09-01 · RANLP 2019 9 · V{\'\i}t Baisa, Ond{\v{r}}ej Herman, Ales Horak

Propaganda of various pressure groups ranging from big economies to ideological blocks is often presented in a form of objective newspaper texts. However, the real objectivity is here shaded with the support of imbalance…

ArticlesPropaganda detection

CUNI Systems in WMT21: Revisiting Backtranslation Techniques for English-Czech NMT

2021-11-01 · WMT (EMNLP) 2021 11 · Petr Gebauer, Ondřej Bojar, Vojtěch Švandelík, Martin Popel

We describe our two NMT systems submitted to the WMT2021 shared task in English-Czech news translation: CUNI-DocTransformer (document-level CUBBITT) and CUNI-Marian-Baselines. We improve the former with a better sentence…

NMTSegmentationSentenceSentence segmentation+1

Text Summarization of Czech News Articles Using Named Entities

2021-04-21 · Petr Marek, Štěpán Müller, Jakub Konrád, Petr Lorenc 외

The foundation for the research of summarization in the Czech language was laid by the work of Straka et al. (2018). They published the SumeCzech, a large Czech news-based summarization dataset, and proposed several base…

Abstractive Text SummarizationArticlesExtractive SummarizationSentence+1