paper-with-me

홈 › Papers

Short text classification with machine learning in the social sciences: The case of climate change on Twitter

2023-10-03 · Karina Shyrokykh, Maksym Girnyk, Lisa Dellmuth

To analyse large numbers of texts, social science researchers are increasingly confronting the challenge of text classification. When manual labeling is not possible and researchers have to find automatized ways to classify texts, computer science provides a useful toolbox of machine-learning methods whose performance remains understudied in the social sciences. In this article, we compare the performance of the most widely used text classifiers by applying them to a typical research scenario in social science research: a relatively small labeled dataset with infrequent occurrence of categories of interest, which is a part of a large unlabeled dataset. As an example case, we look at Twitter communication regarding climate change, a topic of increasing scholarly interest in interdisciplinary social science research. Using a novel dataset including 5,750 tweets from various international organizations regarding the highly ambiguous concept of climate change, we evaluate the performance of methods in automatically classifying tweets based on whether they are about climate change or not. In this context, we highlight two main findings. First, supervised machine-learning methods perform better than state-of-the-art lexicons, in particular as class balance increases. Second, traditional machine-learning methods, such as logistic regression and random forest, perform similarly to sophisticated deep-learning methods, whilst requiring much less training time and computational resources. The results have important implications for the analysis of short texts in social science research.

📄 PDF Abstract BibTeX arXiv:2310.04452

Code (1)

shikarina/short_text_classification 공식 구현 tf

Tasks

text-classificationText Classification

Methods 이 논문이 사용한 방법론

Logistic Regression Logistic Regression, despite its name, is a linear model for classification rather than regression. Logistic regression is also known in the literature as logit regression,…

Similar Papers 제목 키워드 기반

Knowing Your Uncertainty -- On the application of LLM in social sciences

2025-12-05 · Bolun Zhang, Linzhuo Li, Yunqi Chen, Qinlin Zhao 외 arxiv

Large language models (LLMs) are rapidly being integrated into computational social science research, yet their blackboxed training and designed stochastic elements in inference pose unique challenges for scientific inqu…

Machine learning in the social and health sciences

2021-06-20 · Anja K. Leist, Matthias Klee, Jung Hyun Kim, David H. Rehkopf 외

The uptake of machine learning (ML) approaches in the social and health sciences has been rather slow, and research using ML for social and health research questions remains fragmented. This may be due to the separate de…

BIG-bench Machine LearningCausal Inference

SsciBERT: A Pre-trained Language Model for Social Science Texts

2022-06-09 · Si Shen, Jiangfeng Liu, Litao Lin, Ying Huang 외

The academic literature of social sciences records human civilization and studies human social problems. With its large-scale growth, the ways to quickly find existing research on relevant issues have become an urgent de…

Language ModelingLanguage Modellingnamed-entity-recognitionNamed Entity Recognition+1

Aggressive, Repetitive, Intentional, Visible, and Imbalanced: Refining Representations for Cyberbullying Classification

2020-04-04 · Caleb Ziems, Ymir Vigfusson, Fred Morstatter

Cyberbullying is a pervasive problem in online communities. To identify cyberbullying cases in large-scale social networks, content moderators depend on machine learning classifiers for automatic cyberbullying detection.…

General Classification

Using BERT for Qualitative Content Analysis in Psychosocial Online Counseling

2020-11-01 · EMNLP (NLP+CSS) 2020 11 · Philipp Grandeit, Carolyn Haberkern, Maximiliane Lang, Jens Albrecht 외

Qualitative content analysis is a systematic method commonly used in the social sciences to analyze textual data from interviews or online discussions. However, this method usually requires high expertise and manual effo…

BIG-bench Machine Learningtext-classificationText Classification