paper-with-me

홈 › Papers

Benchmarking Machine Learning Robustness in Covid-19 Spike Sequence Classification

2021-09-29 · Sarwan Ali, Bikram Sahoo, Pin-Yu Chen, Murray Patterson

The rapid spread of the COVID-19 pandemic has resulted in an unprecedented amount of sequence data of the SARS-CoV-2 viral genome --- millions of sequences and counting. This amount of data, while being orders of magnitude beyond the capacity of traditional approaches to understanding the diversity, dynamics and evolution of viruses, is nonetheless a rich resource for machine learning (ML) and deep learning (DL) approaches as alternatives for extracting such important information from these data. It is of hence utmost importance to design a framework for testing and benchmarking the robustness of these ML and DL approaches. This paper the first (to our knowledge) to explore such a framework. In this paper, we introduce several ways to perturb SARS-CoV-2 spike protein sequences in ways that mimic the error profiles of common sequencing platforms such as Illumina and PacBio. We show from experiments on a wide array of ML approaches from naive Bayes to logistic regression, that DL approaches are more robust (and accurate) to such adverarial attacks to the input sequences, while $k$-mer based feature vector representations are more robust than the baseline one-hot embedding. Our benchmarking framework may developers of futher ML and DL techniques to properly assess their approaches towards understanding the behaviour of the SARS-CoV-2 virus, or towards avoiding possible future pandemics.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Tasks

BenchmarkingBIG-bench Machine LearningClassification

Similar Papers 제목 키워드 기반

Benchmarking Machine Learning Robustness in Covid-19 Genome Sequence Classification

2022-07-18 · Sarwan Ali, Bikram Sahoo, Alexander Zelikovskiy, Pin-Yu Chen 외

The rapid spread of the COVID-19 pandemic has resulted in an unprecedented amount of sequence data of the SARS-CoV-2 genome -- millions of sequences and counting. This amount of data, while being orders of magnitude beyo…

BenchmarkingBIG-bench Machine LearningClassification

Spike2Vec: An Efficient and Scalable Embedding Approach for COVID-19 Spike Sequences

2021-09-12 · Sarwan Ali, Murray Patterson

With the rapid global spread of COVID-19, more and more data related to this virus is becoming available, including genomic sequence data. The total number of genomic sequences that are publicly available on platforms su…

Variant-driven multi-wave pattern of COVID-19 via a Machine Learning analysis of spike protein mutations

2021-07-21 · Adele de Hoffer, Shahram Vatani, Corentin Cot, Giacomo Cacciapaglia 외

Applying a ML approach to the temporal variability of the Spike protein sequence enables us to identify, classify and track emerging virus variants. Our analysis is unbiased, in the sense that it does not require any pri…

Murmur2Vec: A Hashing Based Solution For Embedding Generation Of COVID-19 Spike Sequences

2025-12-10 · Sarwan Ali, Taslim Murad arxiv

Early detection and characterization of coronavirus disease (COVID-19), caused by SARS-CoV-2, remain critical for effective clinical response and public-health planning. The global availability of large-scale viral seque…

CNN-LSTM Hybrid Model for AI-Driven Prediction of COVID-19 Severity from Spike Sequences and Clinical Data

2025-05-29 · Caio Cheohen, Vinnícius M. S. Gomes, Manuela L. da Silva

The COVID-19 pandemic, caused by SARS-CoV-2, highlighted the critical need for accurate prediction of disease severity to optimize healthcare resource allocation and patient management. The spike protein, which facilitat…

Feature EngineeringRobust classificationseverity prediction