On the Natural Structure of Amino Acid Patterns in Families of Protein Sequences
All known terrestrial proteins are coded as continuous strings of ~20 amino acids. The patterns formed by the repetitions of elements in groups of finite sequences describes the natural architectures of protein families. We present a method to search for patterns and groupings of patterns in protein sequences using a mathematically precise definition for 'repetition', an efficient algorithmic implementation and a robust scoring system with no adjustable parameters. We show that the sequence patterns can be well-separated into disjoint classes according to their recurrence in nested structures. The statistics of pattern occurrences indicate that short repetitions are enough to account for the differences between natural families and randomized groups by more than 10 standard deviations, while patterns shorter than 5 residues are effectively random. A small subset of patterns is sufficient to account for a robust ''familiarity'' definition of arbitrary sets of sequences.
Code (0)
등록된 구현이 없습니다.
Similar Papers 제목 키워드 기반
Computational reconstruction of mitochondria-encoded mammal ancestral proteins
A method based on mapping a symbolic sequence into a set of patterns (strings resulting from the sequence parsing) is proposed as a tool for the reconstruction of ancestral sequences. The set union of patterns comprises …
The Pattern Recognition of Probability Distributions of Amino acids in protein families
A pattern Recognition of a probability distribution of amino acids is obtained for selected families of proteins. The mathematical model is derived from a theory of protein families formation which is derived from applic…
Local sequence-structure relationships in proteins
We seek to understand the interplay between amino acid sequence and local structure in proteins. Are some amino acids unique in their ability to fit harmoniously into certain local structures? What is the role of sequenc…
ClusteringSpecificityInterpretable machine learning of amino acid patterns in proteins: a statistical ensemble approach
Explainable and interpretable unsupervised machine learning helps understand the underlying structure of data. We introduce an ensemble analysis of machine learning models to consolidate their interpretation. Its applica…
Interpretable Machine LearningInferring protein folding mechanisms from natural sequence diversity
Protein sequences serve as a natural record of the evolutionary constraints that shape their functional structures. We show that it is possible to use only sequence information to go beyond predicting native structures a…
DiversityProtein Folding