paper-with-me

홈 › Papers

Czech Dataset for Cross-lingual Subjectivity Classification

2022-04-29 · LREC 2022 6 · Pavel Přibáň, Josef Steinberger

In this paper, we introduce a new Czech subjectivity dataset of 10k manually annotated subjective and objective sentences from movie reviews and descriptions. Our prime motivation is to provide a reliable dataset that can be used with the existing English dataset as a benchmark to test the ability of pre-trained multilingual models to transfer knowledge between Czech and English and vice versa. Two annotators annotated the dataset reaching 0.83 of the Cohen's \k{appa} inter-annotator agreement. To the best of our knowledge, this is the first subjectivity dataset for the Czech language. We also created an additional dataset that consists of 200k automatically labeled sentences. Both datasets are freely available for research purposes. Furthermore, we fine-tune five pre-trained BERT-like models to set a monolingual baseline for the new dataset and we achieve 93.56% of accuracy. We fine-tune models on the existing English dataset for which we obtained results that are on par with the current state-of-the-art results. Finally, we perform zero-shot cross-lingual subjectivity classification between Czech and English to verify the usability of our dataset as the cross-lingual benchmark. We compare and discuss the cross-lingual and monolingual results and the ability of multilingual models to transfer knowledge between languages.

📄 PDF Abstract BibTeX arXiv:2204.13915

Code (2)

pauli31/czech-subjectivity-dataset 공식 구현
pauli31/linear-transformation-4-cs-sa pytorch

Tasks

ClassificationSubjectivity Analysis

Similar Papers 제목 키워드 기반

Are the Multilingual Models Better? Improving Czech Sentiment with Transformers

2021-08-24 · RANLP 2021 9 · Pavel Přibáň, Josef Steinberger

In this paper, we aim at improving Czech sentiment with transformer-based models and their multilingual versions. More concretely, we study the task of polarity detection for the Czech language on three sentiment polarit…

Reading Comprehension in Czech via Machine Translation and Cross-lingual Transfer

2020-07-03 · Kateřina Macková, Milan Straka

Reading comprehension is a well studied task, with huge training datasets in English. This work focuses on building reading comprehension systems for Czech, without requiring any manually annotated Czech training data. F…

Cross-Lingual TransferMachine TranslationReading ComprehensionTranslation

Comparison of Czech Transformers on Text Classification Tasks

2021-07-21 · Jan Lehečka, Jan Švec

In this paper, we present our progress in pre-training monolingual Transformers for Czech and contribute to the research community by releasing our models for public. The need for such models emerged from our effort to e…

Classificationtext-classificationText Classification

Attributivity and Subjectivity in Contemporary Written Czech

2021-12-01 · Quasy (SyntaxFest) 2021 12 · Miroslav Kubát, Radek Čech, Xinying Chen

Linear Transformations for Cross-lingual Sentiment Analysis

2022-09-15 · Pavel Přibáň, Jakub Šmíd, Adam Mištera, Pavel Král

This paper deals with cross-lingual sentiment analysis in Czech, English and French languages. We perform zero-shot cross-lingual classification using five linear transformations combined with LSTM and CNN based classifi…

ClassificationSentiment Analysis