Amino acid frequency and domain features serve well for random forest based classification of thermophilic and mesophilic protein; a case study on serine proteases
Thermostability is an important prerequisite for enzymes employed for industrial applications. Several machine learning based models have thus been formulated for protein classification based on this particular trait. These models have employed features derived from sequences, structures or both resulting in a >93% accuracy based on a 10-fold cross-validation. Besides using various proteins from a wide range of organisms, such studies also rely on hundreds of features. In the present study, an enzyme specific classification model was created using significantly less number of features that provides a similar accuracy of classification for thermophilic and non-thermophilic enzyme serine proteases. For building the classifier, 219 thermophilic and 200 mesophilic bacterial genomes were mined for their respective serine protease sequences. Features were extracted for 800 sequences followed by feature selection. We deployed a random forest based classifier that identified thermophilic and non-thermophilic serine proteases with an accuracy of 95.71%. Knowledge of thermostability along with amino acid positional shifts can be vital for downstream protein engineering techniques. Thus, to emphasize the real time application of the enzyme specific classification model, a web platform has been designed. Combining the sequence data and the classification model, this prototype can allow users to align their query serine protease sequence against the custom database and identify its thermophilic nature.
Code (0)
등록된 구현이 없습니다.
Tasks
Classificationfeature selectionGeneral ClassificationSimilar Papers 제목 키워드 기반
AdaNovo: Adaptive \emph{De Novo} Peptide Sequencing with Conditional Mutual Information
Tandem mass spectrometry has played a pivotal role in advancing proteomics, enabling the analysis of protein composition in biological samples. Despite the development of various deep learning methods for identifying ami…
de novo peptide sequencingNS-Pep: De novo Peptide Design with Non-Standard Amino Acids
Peptide drugs incorporating non-standard amino acids (NSAAs) offer improved binding affinity and improved pharmacological properties. However, existing peptide design methods are limited to standard amino acids, leaving …
Xeno Amino Acids: A look into biochemistry as we don't know it
Would another origin of life resemble Earth's biochemical use of amino acids? Here we review current knowledge at three levels: 1) Could other classes of chemical structure serve as building blocks for biopolymer structu…
Classification of Macromolecule Type Based on Sequences of Amino Acids Using Deep Learning
The classification of amino acids and their sequence analysis plays a vital role in life sciences and is a challenging task. This article uses and compares state-of-the-art deep learning models like convolution neural ne…
General ClassificationPepINVENT: Generative peptide design beyond the natural amino acids
Peptides play a crucial role in the drug design and discovery whether as a therapeutic modality or a delivery agent. Non-natural amino acids (NNAAs) have been used to enhance the peptide properties from binding affinity,…
Drug Design