Document Classification for COVID-19 Literature
The global pandemic has made it more important than ever to quickly and accurately retrieve relevant scientific literature for effective consumption by researchers in a wide range of fields. We provide an analysis of several multi-label document classification models on the LitCovid dataset, a growing collection of 23,000 research papers regarding the novel 2019 coronavirus. We find that pre-trained language models fine-tuned on this dataset outperform all other baselines and that BioBERT surpasses the others by a small margin with micro-F1 and accuracy scores of around 86% and 75% respectively on the test set. We evaluate the data efficiency and generalizability of these models as essential features of any system prepared to deal with an urgent situation like the current health crisis. Finally, we explore 50 errors made by the best performing models on LitCovid documents and find that they often (1) correlate certain labels too closely together and (2) fail to focus on discriminative sections of the articles; both of which are important issues to address in future work. Both data and code are available on GitHub.
Code (1)
Tasks
ArticlesClassificationDocument ClassificationGeneral ClassificationMethods 이 논문이 사용한 방법론
Similar Papers 제목 키워드 기반
Classification of COVID19 tweets using Machine Learning Approaches
The reported work is a description of our participation in the “Classification of COVID19 tweets containing symptoms” shared task, organized by the “Social Media Mining for Health Applications (SMM4H)” workshop. The lite…
BIG-bench Machine LearningClassificationNavigating the landscape of COVID-19 research through literature analysis: A bird's eye view
Timely access to accurate scientific literature in the battle with the ongoing COVID-19 pandemic is critical. This unprecedented public health risk has motivated research towards understanding the disease in general, ide…
ArticlesClusteringnamed-entity-recognitionNamed Entity Recognition+2COV19IR : COVID-19 Domain Literature Information Retrieval
Increasing number of COVID-19 research literatures cause new challenges in effective literature screening and COVID-19 domain knowledge aware Information Retrieval. To tackle the challenges, we demonstrate two tasks alon…
Information RetrievalQuestion AnsweringRetrievalS_Covid: An Engine to Explore COVID-19 Scientific Literature
This paper introduces S_Covid, an end-to-end unsupervised learning based question-answering engine for exploring COVID-19 scientific literature collections. S_Covid enables documents exploration for finding relevant rese…
Information RetrievalQuestion AnsweringRetrievalPrioritization of COVID-19-related literature via unsupervised keyphrase extraction and document representation learning
The COVID-19 pandemic triggered a wave of novel scientific literature that is impossible to inspect and study in a reasonable time frame manually. Current machine learning methods offer to project such body of literature…
Keyphrase ExtractionRepresentation Learning