paper-with-me

Papers

IRLCov19: A Large COVID-19 Multilingual Twitter Dataset of Indian Regional Languages

2021-07-26 · Deepak Uniyal, Amit Agarwal

Emerged in Wuhan city of China in December 2019, COVID-19 continues to spread rapidly across the world despite authorities having made available a number of vaccines. While the coronavirus has been around for a significant period of time, people and authorities still feel the need for awareness due to the mutating nature of the virus and therefore varying symptoms and prevention strategies. People and authorities resort to social media platforms the most to share awareness information and voice out their opinions due to their massive outreach in spreading the word in practically no time. People use a number of languages to communicate over social media platforms based on their familiarity, language outreach, and availability on social media platforms. The entire world has been hit by the coronavirus and India is the second worst-hit country in terms of the number of active coronavirus cases. India, being a multilingual country, offers a great opportunity to study the outreach of various languages that have been actively used across social media platforms. In this study, we aim to study the dataset related to COVID-19 collected in the period between February 2020 to July 2020 specifically for regional languages in India. This could be helpful for the Government of India, various state governments, NGOs, researchers, and policymakers in studying different issues related to the pandemic. We found that English has been the mode of communication in over 64% of tweets while as many as twelve regional languages in India account for approximately 4.77% of tweets.

📄 PDF Abstract BibTeX arXiv:2107.12360

Code (1)

deepakuniyaliit/Covid19IRLTDataset 공식 구현

Similar Papers 제목 키워드 기반

NAIST COVID: Multilingual COVID-19 Twitter and Weibo Dataset

2020-04-17 · Zhiwei Gao, Shuntaro Yada, Shoko Wakamiya, Eiji Aramaki

Since the outbreak of coronavirus disease 2019 (COVID-19) in the late 2019, it has affected over 200 countries and billions of people worldwide. This has affected the social life of people owing to enforcements, such as …

COVID-19 Misinformation on Twitter: Multilingual Analysis

2021-01-06 · Raj Ratn Pranesh, Mehrdad Farokhenajd, Ambesh Shekhar, Genoveva Vargas-Solar

In the current scenario of the coronavirus disease pandemic (COVID-19), the Internet has become an important source of health information for users worldwide. During pandemic situations, myths, sensationalism, rumours an…

MisinformationRumour Detection

CMTA: COVID-19 Misinformation Multilingual Analysis on Twitter

2021-08-01 · ACL 2021 5 · Raj Pranesh, Mehrdad Farokhenajd, Ambesh Shekhar, Genoveva Vargas-Solar

The internet has actually come to be an essential resource of health knowledge for individuals around the world in the present situation of the coronavirus condition pandemic(COVID-19). During pandemic situations, myths,…

MisinformationRumour DetectionTransfer Learning

COVID-Twitter-BERT: A Natural Language Processing Model to Analyse COVID-19 Content on Twitter

2020-05-15 · Martin Müller, Marcel Salathé, Per E Kummervold

In this work, we release COVID-Twitter-BERT (CT-BERT), a transformer-based model, pretrained on a large corpus of Twitter messages on the topic of COVID-19. Our model shows a 10-30% marginal improvement compared to its b…

ClassificationGeneral ClassificationQuestion Answering

Cross-language sentiment analysis of European Twitter messages during the COVID-19 pandemic

2020-07-01 · ACL 2020 7 · Anna Kruspe, Matthias H{\"a}berle, Iona Kuhn, Xiao Xiang Zhu

In this paper, we analyze Twitter messages (tweets) collected during the first months of the COVID-19 pandemic in Europe with regard to their sentiment. This is implemented with a neural network for sentiment analysis us…

SentenceSentence EmbeddingsSentiment Analysis