paper-with-me

홈 › Papers

Small data problems in political research: a critical replication study

2021-09-27 · Hugo de Vos, Suzan Verberne

In an often-cited 2019 paper on the use of machine learning in political research, Anastasopoulos & Whitford (A&W) propose a text classification method for tweets related to organizational reputation. The aim of their paper was to provide a 'guide to practice' for public administration scholars and practitioners on the use of machine learning. In the current paper we follow up on that work with a replication of A&W's experiments and additional analyses on model stability and the effects of preprocessing, both in relation to the small data size. We show that (1) the small data causes the classification model to be highly sensitive to variations in the random train-test split, and that (2) the applied preprocessing causes the data to be extremely sparse, with the majority of items in the data having at most two non-zero lexical features. With additional experiments in which we vary the steps of the preprocessing pipeline, we show that the small data size keeps causing problems, irrespective of the preprocessing choices. Based on our findings, we argue that A&W's conclusions regarding the automated classification of organizational reputation tweets -- either substantive or methodological -- can not be maintained and require a larger data set for training and more careful validation.

📄 PDF Abstract BibTeX arXiv:2109.12911

Code (0)

등록된 구현이 없습니다.

Tasks

text-classificationText Classification

Similar Papers 제목 키워드 기반

Fairness Research For Machine Learning Should Integrate Societal Considerations

2025-06-14 · Yijun Bian, Lei You

Enhancing fairness in machine learning (ML) systems is increasingly important nowadays. While current research focuses on assistant tools for ML pipelines to promote fairness within them, we argue that: 1) The significan…

Fairness

Computational Analysis of Political Texts: Bridging Research Efforts Across Communities

2019-07-01 · ACL 2019 7 · Goran Glava{\v{s}}, Federico Nanni, Simone Paolo Ponzetto

In the last twenty years, political scientists started adopting and developing natural language processing (NLP) methods more actively in order to exploit text as an additional source of data in their analyses. Over the …

Stance Detection

Creating Japanese Political Corpus from Local Assembly Minutes of 47 prefectures

2016-12-01 · WS 2016 12 · Yasutomo Kimura, Keiichi Takamaru, Takuma Tanaka, Akio Kobayashi 외

This paper describes a Japanese political corpus created for interdisciplinary political research. The corpus contains the local assembly minutes of 47 prefectures from April 2011 to March 2015. This four-year period coi…

High Risk of Political Bias in Black Box Emotion Inference Models

2024-07-18 · Hubert Plisiecki, Paweł Lenartowicz, Maria Flakus, Artur Pokropek

This paper investigates the presence of political bias in emotion inference models used for sentiment analysis (SA) in social science research. Machine learning models often reflect biases in their training data, impacti…

Sentiment Analysis

Inducing Political Bias Allows Language Models Anticipate Partisan Reactions to Controversies

2023-11-16 · Zihao He, Siyi Guo, Ashwin Rao, Kristina Lerman

Social media platforms are rife with politically charged discussions. Therefore, accurately deciphering and predicting partisan biases using Large Language Models (LLMs) is increasingly critical. In this study, we addres…

Stance Detection