paper-with-me

홈 › Papers

COVID-19 Mythbusters in World Languages

2022-06-01 · LREC 2022 6 · Mana Ashida, Jin-Dong Kim, Lee Seunghun

This paper introduces a multi-lingual database containing translated texts of COVID-19 mythbusters. The database has translations into 115 languages as well as the original English texts, of which the original texts are published by World Health Organization (WHO). This paper then presents preliminary analyses on latin-alphabet-based texts to see the potential of the database as a resource for multilingual linguistic analyses. The analyses on latin-alphabet-based texts gave interesting insights into the resource. While the amount of translated texts in each language was small, character bi-grams with normalization (lowercasing and removal of diacritics) was turned out to be an effective proxy for measuring the similarity of the languages, and the affinity ranking of language pairs could be obtained. Additionally, the hierarchical clustering analysis is performed using the character bigram overlap ratio of every possible pair of languages. The result shows the cluster of Germanic languages, Romance languages, and Southern Bantu languages. In sum, the multilingual database not only offers fixed set of materials in numerous languages, but also serves as a preliminary tool to identify the language family using text-based similarity measure of bigram overlap ratio.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

TICO-19: the Translation Initiative for Covid-19

2020-07-03 · EMNLP (NLP-COVID19) 2020 12 · Antonios Anastasopoulos, Alessandro Cattelan, Zi-Yi Dou, Marcello Federico 외

The COVID-19 pandemic is the worst pandemic to strike the world in over a century. Crucial to stemming the tide of the SARS-CoV-2 virus is communicating to vulnerable populations the means by which they can protect thems…

Translation

A System for Worldwide COVID-19 Information Aggregation

2020-07-28 · EMNLP (NLP-COVID19) 2020 12 · Akiko Aizawa, Frederic Bergeron, Junjie Chen, Fei Cheng 외

The global pandemic of COVID-19 has made the public pay close attention to related news, covering various domains, such as sanitation, treatment, and effects on education. Meanwhile, the COVID-19 condition is very differ…

ArticlesMachine TranslationTranslation

IRLCov19: A Large COVID-19 Multilingual Twitter Dataset of Indian Regional Languages

2021-07-26 · Deepak Uniyal, Amit Agarwal

Emerged in Wuhan city of China in December 2019, COVID-19 continues to spread rapidly across the world despite authorities having made available a number of vaccines. While the coronavirus has been around for a significa…

CovTiNet: Covid text identification network using attention-based positional embedding feature fusion

2023-03-14 · Neural Computing and Applications 2023 3 · Md. Rajib Hossain, Mohammed Moshiul Hoque, Nazmul Siddique & Iqbal H. Sarker

Covid text identification (CTI) is a crucial research concern in natural language processing (NLP). Social and electronic media are simultaneously adding a large volume of Covid-affiliated text on the World Wide Web due …

Misinformation

SenWave: Monitoring the Global Sentiments under the COVID-19 Pandemic

2020-06-18 · Qiang Yang, Hind Alamro, Somayah Albaradei, Adil Salhi 외

Since the first alert launched by the World Health Organization (5 January, 2020), COVID-19 has been spreading out to over 180 countries and territories. As of June 18, 2020, in total, there are now over 8,400,000 cases …

Language ModellingSentiment Analysis