Application of Markov Structure of Genomes to Outlier Identification and Read Classification
In this paper we apply the structure of genomes as second-order Markov processes specified by the distributions of successive triplets of bases to two bioinformatics problems: identification of outliers in genome databases and read classification in metagenomics, using real coronavirus and adenovirus data.
Code (0)
등록된 구현이 없습니다.
Similar Papers 제목 키워드 기반
Structured and Unstructured Outlier Identification for Robust PCA: A Non iterative, Parameter free Algorithm
Robust PCA, the problem of PCA in the presence of outliers has been extensively investigated in the last few years. Here we focus on Robust PCA in the outlier model where each column of the data matrix is either an inlie…
Outlier robust system identification: a Bayesian kernel-based approach
In this paper, we propose an outlier-robust regularized kernel-based method for linear system identification. The unknown impulse response is modeled as a zero-mean Gaussian process whose covariance (kernel) is given by …
A Markovian genomic concatenation model guided by persymmetric matrices
The aim of this work is to provide a rigorous mathematical analysis of a stochastic concatenation model presented by Sobottka and Hart (2011) which allows approximation of the first-order stochastic structure in bacteria…
LinearSankoff: Linear-time Simultaneous Folding and Alignment of RNA Homologs
The classical Sankoff algorithm for the simultaneous folding and alignment of homologous RNA sequences is highly influential, but it suffers from two major limitations in efficiency and modeling power. First, it takes $O…
Identification of repeats in DNA sequences using nucleotide distribution uniformity
Repetitive elements are important in genomic structures, functions and regulations, yet effective methods in precisely identifying repetitive elements in DNA sequences are not fully accessible, and the relationship betwe…