POLITICS: Pretraining with Same-story Article Comparison for Ideology Prediction and Stance Detection
Ideology is at the core of political science research. Yet, there still does not exist general-purpose tools to characterize and predict ideology across different genres of text. To this end, we study Pretrained Language Models using novel ideology-driven pretraining objectives that rely on the comparison of articles on the same story written by media of different ideologies. We further collect a large-scale dataset, consisting of more than 3.6M political news articles, for pretraining. Our model POLITICS outperforms strong baselines and the previous state-of-the-art models on ideology prediction and stance detection tasks. Further analyses show that POLITICS is especially good at understanding long or formally written texts, and is also robust in few-shot learning scenarios.
Code (2)
Tasks
ArticlesFew-Shot LearningStance DetectionSimilar Papers 제목 키워드 기반
POLITICS: Pretraining with Same-story Article Comparison for Ideology Prediction and Stance Detection
Ideology is at the core of political science. Yet, there still does not exist general-purpose tools that can characterize and predict ideology across different genres of text. To this end, we study the training of PLMs …
ArticlesFew-Shot LearningStance DetectionBREAKING! Presenting Fake News Corpus for Automated Fact Checking
Popular fake news articles spread faster than mainstream articles on the same topic which renders manual fact checking inefficient. At the same time, creating tools for automatic detection is as challenging due to lack o…
ArticlesFact CheckingFake News DetectionAll Things Considered: Detecting Partisan Events from News Media with Cross-Article Comparison
Public opinion is shaped by the information news media provide, and that information in turn may be shaped by the ideological preferences of media outlets. But while much attention has been devoted to media bias via over…
AllArticlesCultural Topic Modelling over Novel Wikipedia Corpora for South-Slavic Languages
There is a shortage of high-quality corpora for South-Slavic languages. Such corpora are useful to computer scientists and researchers in social sciences and humanities alike, focusing on numerous linguistic, content ana…
Cultural Vocal Bursts Intensity PredictionHistoryComparator: Interactive Across-Time Comparison in Document Archives
Recent years have witnessed significant increase in the number of large scale digital collections of archival documents such as news articles, books, etc. Typically, users access these collections through searching or br…
ArticlesClusteringDecision MakingNamed Entity Recognition (NER)