How COVID-19 Is Changing Our Language : Detecting Semantic Shift in Twitter Word Embeddings
Words are malleable objects, influenced by events that are reflected in written texts. Situated in the global outbreak of COVID-19, our research aims at detecting semantic shifts in social media language triggered by the health crisis. With COVID-19 related big data extracted from Twitter, we train separate word embedding models for different time periods after the outbreak. We employ an alignment-based approach to compare these embeddings with a general-purpose Twitter embedding unrelated to COVID-19. We also compare our trained embeddings among them to observe diachronic evolution. Carrying out case studies on a set of words chosen by topic detection, we verify that our alignment approach is valid. Finally, we quantify the size of global semantic shift by a stability measure based on back-and-forth rotational alignment.
Code (0)
등록된 구현이 없습니다.
Tasks
validWord EmbeddingsSimilar Papers 제목 키워드 기반
The Problem of Semantic Shift in Longitudinal Monitoring of Social Media: A Case Study on Mental Health During the COVID-19 Pandemic
Social media allows researchers to track societal and cultural changes over time based on language analysis tools. Many of these tools rely on statistical algorithms which need to be tuned to specific types of language. …
Fine-Tuning Deteriorates General Textual Out-of-Distribution Detection by Distorting Task-Agnostic Features
Detecting out-of-distribution (OOD) inputs is crucial for the safe deployment of natural language processing (NLP) models. Though existing methods, especially those based on the statistics in the feature space of fine-tu…
Out-of-Distribution DetectionOut of Distribution (OOD) DetectionA Comparative Study of Hybrid Models in Health Misinformation Text Classification
This study evaluates the effectiveness of machine learning (ML) and deep learning (DL) models in detecting COVID-19-related misinformation on online social networks (OSNs), aiming to develop more effective tools for coun…
LemmatizationMisinformationtext-classificationText ClassificationDetecting COVID-19 Conspiracy Theories with Transformers and TF-IDF
The sharing of fake news and conspiracy theories on social media has wide-spread negative effects. By designing and applying different machine learning models, researchers have made progress in detecting fake news from t…
Common Sense ReasoningFake News DetectionUsing neural topic models to track context shifts of words: a case study of COVID-related terms before and after the lockdown in April 2020
This paper explores lexical meaning changes in a new dataset, which includes tweets from before and after the COVID-related lockdown in April 2020. We use this dataset to evaluate traditional and more recent unsupervised…
Language ModelingLanguage ModellingTopic Models