paper-with-me

Papers

Stanceosaurus: Classifying Stance Towards Multilingual Misinformation

2022-10-28 · Jonathan Zheng, Ashutosh Baheti, Tarek Naous, Wei Xu, Alan Ritter

We present Stanceosaurus, a new corpus of 28,033 tweets in English, Hindi, and Arabic annotated with stance towards 251 misinformation claims. As far as we are aware, it is the largest corpus annotated with stance towards misinformation claims. The claims in Stanceosaurus originate from 15 fact-checking sources that cover diverse geographical regions and cultures. Unlike existing stance datasets, we introduce a more fine-grained 5-class labeling strategy with additional subcategories to distinguish implicit stance. Pre-trained transformer-based stance classifiers that are fine-tuned on our corpus show good generalization on unseen claims and regional claims from countries outside the training data. Cross-lingual experiments demonstrate Stanceosaurus' capability of training multi-lingual models, achieving 53.1 F1 on Hindi and 50.4 F1 on Arabic without any target-language fine-tuning. Finally, we show how a domain adaptation method can be used to improve performance on Stanceosaurus using additional RumourEval-2019 data. We make Stanceosaurus publicly available to the research community and hope it will encourage further work on misinformation identification across languages and cultures.

📄 PDF Abstract BibTeX arXiv:2210.15954

Code (0)

등록된 구현이 없습니다.

Tasks

Domain AdaptationFact CheckingMisinformation

Similar Papers 제목 키워드 기반

Stanceosaurus 2.0: Classifying Stance Towards Russian and Spanish Misinformation

2024-02-06 · Anton Lavrouk, Ian Ligon, Tarek Naous, Jonathan Zheng 외

The Stanceosaurus corpus (Zheng et al., 2022) was designed to provide high-quality, annotated, 5-way stance data extracted from Twitter, suitable for analyzing cross-cultural and cross-lingual misinformation. In the Stan…

Cross-Lingual TransferMisinformationStance ClassificationZero-Shot Cross-Lingual Transfer

Looking for COVID-19 misinformation in multilingual social media texts

2021-05-03 · Raj Ratn Pranesh, Mehrdad Farokhnejad, Ambesh Shekhar, Genoveva Vargas-Solar

This paper presents the Multilingual COVID-19 Analysis Method (CMTA) for detecting and observing the spread of misinformation about this disease within texts. CMTA proposes a data science (DS) pipeline that applies machi…

Misinformation

Lost in Translation -- Multilingual Misinformation and its Evolution

2023-10-27 · Dorian Quelle, Calvin Cheng, Alexandre Bovet, Scott A. Hale

Misinformation and disinformation are growing threats in the digital age, spreading rapidly across languages and borders. This paper investigates the prevalence and dynamics of multilingual misinformation through an anal…

Fact CheckingMisinformationSentenceSentence Embeddings+1

CMTA: COVID-19 Misinformation Multilingual Analysis on Twitter

2021-08-01 · ACL 2021 5 · Raj Pranesh, Mehrdad Farokhenajd, Ambesh Shekhar, Genoveva Vargas-Solar

The internet has actually come to be an essential resource of health knowledge for individuals around the world in the present situation of the coronavirus condition pandemic(COVID-19). During pandemic situations, myths,…

MisinformationRumour DetectionTransfer Learning

Emotion Detection for Misinformation: A Review

2023-11-01 · Zhiwei Liu, Tianlin Zhang, Kailai Yang, Paul Thompson 외

With the advent of social media, an increasing number of netizens are sharing and reading posts and news online. However, the huge volumes of misinformation (e.g., fake news and rumors) that flood the internet can advers…

Fake News DetectionMisinformation