paper-with-me

Papers

MuMiN: A Large-Scale Multilingual Multimodal Fact-Checked Misinformation Social Network Dataset

2022-02-23 · Dan Saattrup Nielsen, Ryan McConville

Misinformation is becoming increasingly prevalent on social media and in news articles. It has become so widespread that we require algorithmic assistance utilising machine learning to detect such content. Training these machine learning models require datasets of sufficient scale, diversity and quality. However, datasets in the field of automatic misinformation detection are predominantly monolingual, include a limited amount of modalities and are not of sufficient scale and quality. Addressing this, we develop a data collection and linking system (MuMiN-trawl), to build a public misinformation graph dataset (MuMiN), containing rich social media data (tweets, replies, users, images, articles, hashtags) spanning 21 million tweets belonging to 26 thousand Twitter threads, each of which have been semantically linked to 13 thousand fact-checked claims across dozens of topics, events and domains, in 41 different languages, spanning more than a decade. The dataset is made available as a heterogeneous graph via a Python package (mumin). We provide baseline results for two node classification tasks related to the veracity of a claim involving social media, and demonstrate that these are challenging tasks, with the highest macro-average F1-score being 62.55% and 61.45% for the two tasks, respectively. The MuMiN ecosystem is available at https://mumin-dataset.github.io/, including the data, documentation, tutorials and leaderboards.

📄 PDF Abstract BibTeX arXiv:2202.11684

Code (3)

MuMiN-dataset/mumin-build 공식 구현 pytorch
MuMiN-dataset/mumin-baseline pytorch
MuMiN-dataset/mumin-trawl pytorch

Tasks

ArticlesMisinformationNode Classification

Similar Papers 제목 키워드 기반

Automating Claim Construction in Patent Applications: The CMUmine Dataset

2021-11-01 · EMNLP (NLLP) 2021 11 · Ozan Tonguz, Yiwei Qin, Yimeng Gu, Hyun Hannah Moon

Intellectual Property (IP) in the form of issued patents is a critical and very desirable element of innovation in high-tech. In this position paper, we explore the possibility of automating the legal task of Claim Const…

Position

mOSCAR: A Large-scale Multilingual and Multimodal Document-level Corpus

2024-06-13 · Matthieu Futeral, Armel Zebaze, Pedro Ortiz Suarez, Julien Abadji 외

Multimodal Large Language Models (mLLMs) are trained on a large amount of text-image data. While most mLLMs are trained on caption-like data only, Alayrac et al. [2022] showed that additionally training them on interleav…

Few-Shot LearningIn-Context Learning

Semi-automated Fact-checking in Portuguese: Corpora Enrichment using Retrieval with Claim extraction

2025-07-19 · Juliana Resplande Sant'anna Gomes, Arlindo Rodrigues Galvão Filho arxiv

The accelerated dissemination of disinformation often outpaces the capacity for manual fact-checking, highlighting the urgent need for Semi-Automated Fact-Checking (SAFC) systems. Within the Portuguese language context, …

Multilingual, Multimodal Pipeline for Creating Authentic and Structured Fact-Checked Claim Dataset

2026-01-12 · Z. Melce Hüsünbeyi, Virginie Mouilleron, Leonie Uhling, Daniel Foppe 외 arxiv

The rapid proliferation of misinformation across online platforms underscores the urgent need for robust, up-to-date, explainable, and multilingual fact-checking resources. However, existing datasets are limited in scope…

Added Toxicity Mitigation at Inference Time for Multimodal and Massively Multilingual Translation

2023-11-11 · Marta R. Costa-jussà, David Dale, Maha Elbayad, Bokai Yu

Added toxicity in the context of translation refers to the fact of producing a translation output with more toxicity than there exists in the input. In this paper, we present MinTox which is a novel pipeline to identify …

Machine TranslationTranslation