Navigating Homogeneous Paths through Amyloidogenic and Non-Amyloidogenic Hexapeptides
Hexapeptides are increasingly applied as model systems for studying the amyloidogenecity properties of oligo- and polypeptides. It is possible to construct 64 million different hexapeptides from the twenty proteinogenic amino acid residues. Today's experimental amyloid databases contain only a fraction of these annotated hexapeptides. For labeling all the possible hexapeptides as "amyloidogenic" or "non-amyloidogenic" there exist several computational predictors with good accuracies. It may be of interest to define and study a simple graph structure on the 64 million hexapeptides as nodes when two hexapeptides are connected by an edge if they differ by only a single residue. For example, in this graph, HIKKLM is connected to AIKKLM, or HIKKNM, or HIKKLC, but it is not connected with an edge to VVKKLM or HIKNPM. In the present contribution, we consider our previously published artificial intelligence-based tool, the Budapest Amyloid Predictor (BAP for short), and demonstrate a spectacular property of this predictor in the graph defined above. We show that for any two hexapeptides predicted to be "amyloidogenic" by the BAP predictor, there exists an easily constructible path of length at most 6 that passes through neighboring hexapeptides all predicted to be "amyloidogenic" by BAP. For example, the predicted amyloidogenic ILVWIW and FWLCYL hexapeptides can be connected through the length-6 path ILVWIW-IWVWIW-IWVCIW-IWVCIL-FWVCIL-FWLCIL-FWLCYL in such a way that the neighbors differ in exactly one residue, and all hexapeptides on the path are predicted to be amyloidogenic by BAP. The symmetric statement also holds for non-amyloidogenic hexapeptides. It is noted that the mentioned property of the Budapest Amyloid Predictor \url{https://pitgroup.org/bap} is not proprietary; it is also true for any linear Support Vector Machine (SVM)-based predictors.
Code (0)
등록된 구현이 없습니다.
Similar Papers 제목 키워드 기반
Succinct Amyloid and Non-Amyloid Patterns in Hexapeptides
Hexapeptides are widely applied as a model system for studying amyloid-forming properties of polypeptides, including proteins. Recently, large experimental databases have become publicly available with amyloidogenic labe…
Deep Learning Model for Amyloidogenicity Prediction using a Pre-trained Protein LLM
The prediction of amyloidogenicity in peptides and proteins remains a focal point of ongoing bioinformatics. The crucial step in this field is to apply advanced computational methodologies. Many recent approaches to pred…
Opening Amyloid-Windows to the Secondary Structure of Proteins: The Amyloidogenecity Increases Tenfold Inside Beta-Sheets
Methods from artificial intelligence (AI), in general, and machine learning, in particular, have kept conquering new territories in numerous areas of science. Most of the applications of these techniques are restricted t…
Citrate stabilized gold nanoparticles interfere with amyloid fibril formation: D76N and {\Delta}N6 \b{eta}2-microglobulin variants
Protein aggregation including the formation of dimers and multimers in solution, underlies an array of human diseases such as systemic amyloidosis which is a fatal disease caused by misfolding of native globular proteins…
The structure of N184K amyloidogenic variant of gelsolin highlights the role of the H-bond network for protein stability and aggregation properties
Mutations in the gelsolin protein are responsible for a rare conformational disease known as AGel amyloidosis. Four of these mutations are hosted by the second domain of the protein (G2): D187N/Y, G167R and N184K. The im…