Telling apart <I>Felidae</I> and <I>Ursidae</I> from the distribution of nucleotides in mitochondrial DNA
Rank--frequency distributions of nucleotide sequences in mitochondrial DNA are defined in a way analogous to the linguistic approach, with the highest-frequent nucleobase serving as a whitespace. For such sequences, entropy and mean length are calculated. These parameters are shown to discriminate the species of the <I>Felidae</I> (cats) and <I>Ursidae</I> (bears) families. From purely numerical values we are able to see in particular that giant pandas are bears while koalas are not. The observed linear relation between the parameters is explained using a simple probabilistic model. The approach based on the nonadditive generalization of the Bose-distribution is used to analyze the frequency spectra of the nucleotide sequences. In this case, the separation of families is not very sharp. Nevertheless, the distributions for <I>Felidae</I> have on average longer tails comparing to <I>Ursidae</I>.
Code (0)
등록된 구현이 없습니다.
Similar Papers 제목 키워드 기반
On the verge of life: Distribution of nucleotide sequences in viral RNAs
The aim of the study is to analyze viruses using parameters obtained from distributions of nucleotide sequences in the viral RNA. Seeking for the input data homogeneity, we analyze single-stranded RNA viruses only. Two a…
General ClassificationDependent Multinomial Models Made Easy: Stick Breaking with the Pólya-Gamma Augmentation
Many practical modeling problems involve discrete data that are best represented as draws from multinomial or categorical distributions. For example, nucleotides in a DNA sequence, children's names in a given state and y…
Bayesian InferencePositionAdd a SideNet to your MainNet
As the performance and popularity of deep neural networks has increased, so too has their computational cost. There are many effective techniques for reducing a network's computational footprint (quantisation, pruning, k…
General ClassificationKnowledge Distillationtext-classificationText ClassificationDependent Multinomial Models Made Easy: Stick-Breaking with the Polya-gamma Augmentation
Many practical modeling problems involve discrete data that are best represented as draws from multinomial or categorical distributions. For example, nucleotides in a DNA sequence, children's names in a given state and y…
Bayesian InferenceFeature extraction from complex networks: A case of study in genomic sequences classification
This work presents a new approach for classification of genomic sequences from measurements of complex networks and information theory. For this, it is considered the nucleotides, dinucleotides and trinucleotides of a ge…
ClassificationGeneral Classification