Auditing and Robustifying COVID-19 Misinformation Datasets via Anticontent Sampling
This paper makes two key contributions. First, it argues that highly specialized rare content classifiers trained on small data typically have limited exposure to the richness and topical diversity of the negative class (dubbed anticontent) as observed in the wild. As a result, these classifiers' strong performance observed on the test set may not translate into real-world settings. In the context of COVID-19 misinformation detection, we conduct an in-the-wild audit of multiple datasets and demonstrate that models trained with several prominently cited recent datasets are vulnerable to anticontent when evaluated in the wild. Second, we present a novel active learning pipeline that requires zero manual annotation and iteratively augments the training data with challenging anticontent, robustifying these classifiers.
Code (0)
등록된 구현이 없습니다.
Tasks
Active LearningDiversityMisinformationSimilar Papers 제목 키워드 기반
The COVMis-Stance dataset: Stance Detection on Twitter for COVID-19 Misinformation
During the COVID-19 pandemic, large amounts of COVID-19 misinformation are spreading on social media. We are interested in the stance of Twitter users towards COVID-19 misinformation. However, due to the relative recent …
MisinformationStance DetectionTesting the Generalization of Neural Language Models for COVID-19 Misinformation Detection
A drastic rise in potentially life-threatening misinformation has been a by-product of the COVID-19 pandemic. Computational support to identify false information within the massive body of data on the topic is crucial to…
ArticlesMisinformationCOVIDLies: Detecting COVID-19 Misinformation on Social Media
The ongoing pandemic has heightened the need for developing tools to flag COVID-19-related misinformation on the internet, specifically on social media such as Twitter. However, due to novel language and the rapid change…
MisconceptionsMisinformationRetrievalStance DetectionAMIR: Automated MisInformation Rebuttal -- A COVID-19 Vaccination Datasets based Recommendation System
Misinformation has emerged as a major societal threat in recent years in general; specifically in the context of the COVID-19 pandemic, it has wrecked havoc, for instance, by fuelling vaccine hesitancy. Cost-effective, s…
ArticlesMisinformationNot cool, calm or collected: Using emotional language to detect COVID-19 misinformation
COVID-19 misinformation on social media platforms such as twitter is a threat to effective pandemic management. Prior works on tweet COVID-19 misinformation negates the role of semantic features common to twitter such as…
ManagementMisinformation