paper-with-me

홈 › Papers

Mega-COV: A Billion-Scale Dataset of 100+ Languages for COVID-19

2020-05-02 · EACL 2021 2 · Muhammad Abdul-Mageed, AbdelRahim Elmadany, El Moatez Billah Nagoudi, Dinesh Pabbi, Kunal Verma, Rannie Lin

We describe Mega-COV, a billion-scale dataset from Twitter for studying COVID-19. The dataset is diverse (covers 268 countries), longitudinal (goes as back as 2007), multilingual (comes in 100+ languages), and has a significant number of location-tagged tweets (~169M tweets). We release tweet IDs from the dataset. We also develop and release two powerful models, one for identifying whether or not a tweet is related to the pandemic (best F1=97%) and another for detecting misinformation about COVID-19 (best F1=92%). A human annotation study reveals the utility of our models on a subset of Mega-COV. Our data and models can be useful for studying a wide host of phenomena related to the pandemic. Mega-COV and our models are publicly available.

📄 PDF Abstract BibTeX arXiv:2005.06012

Code (1)

UBC-NLP/megacov 공식 구현

Tasks

Misinformation

Similar Papers 제목 키워드 기반

COVID-19-related Nepali Tweets Classification in a Low Resource Setting

2022-10-11 · SMM4H (COLING) 2022 10 · Rabin Adhikari, Safal Thapaliya, Nirajan Basnet, Samip Poudel 외

Billions of people across the globe have been using social media platforms in their local languages to voice their opinions about the various topics related to the COVID-19 pandemic. Several organizations, including the …

Megalodon: Efficient LLM Pretraining and Inference with Unlimited Context Length

2024-04-12 · Xuezhe Ma, Xiaomeng Yang, Wenhan Xiong, Beidi Chen 외

The quadratic complexity and weak length extrapolation of Transformers limits their ability to scale to long sequences, and while sub-quadratic solutions like linear attention and state space models exist, they empirical…

State Space Models

SenWave: A Fine-Grained Multi-Language Sentiment Analysis Dataset Sourced from COVID-19 Tweets

2025-10-09 · Qiang Yang, Xiuying Chen, Changsheng Ma, Rui Yin 외 arxiv

The global impact of the COVID-19 pandemic has highlighted the need for a comprehensive understanding of public sentiment and reactions. Despite the availability of numerous public datasets on COVID-19, some reaching vol…

Sentiment Analysis

Regular omega-Languages with an Informative Right Congruence

2018-09-10 · Dana Angluin, Dana Fisman

A regular language is almost fully characterized by its right congruence relation. Indeed, a regular language can always be recognized by a DFA isomorphic to the automaton corresponding to its right congruence, hencefort…

mGPT: Few-Shot Learners Go Multilingual

2022-04-15 · Oleh Shliazhko, Alena Fenogenova, Maria Tikhonova, Vladislav Mikhailov 외

Recent studies report that autoregressive language models can successfully solve many NLP tasks via zero- and few-shot learning paradigms, which opens up new possibilities for using the pre-trained language models. This …

Cross-Lingual Natural Language InferenceCross-Lingual Paraphrase IdentificationCross-Lingual TransferFew-Shot Learning+4