paper-with-me

홈 › Papers

Amino acid frequency and domain features serve well for random forest based classification of thermophilic and mesophilic protein; a case study on serine proteases

2021-03-05 · Jithin S. Sunny, Lilly M. Saleena

Thermostability is an important prerequisite for enzymes employed for industrial applications. Several machine learning based models have thus been formulated for protein classification based on this particular trait. These models have employed features derived from sequences, structures or both resulting in a >93% accuracy based on a 10-fold cross-validation. Besides using various proteins from a wide range of organisms, such studies also rely on hundreds of features. In the present study, an enzyme specific classification model was created using significantly less number of features that provides a similar accuracy of classification for thermophilic and non-thermophilic enzyme serine proteases. For building the classifier, 219 thermophilic and 200 mesophilic bacterial genomes were mined for their respective serine protease sequences. Features were extracted for 800 sequences followed by feature selection. We deployed a random forest based classifier that identified thermophilic and non-thermophilic serine proteases with an accuracy of 95.71%. Knowledge of thermostability along with amino acid positional shifts can be vital for downstream protein engineering techniques. Thus, to emphasize the real time application of the enzyme specific classification model, a web platform has been designed. Combining the sequence data and the classification model, this prototype can allow users to align their query serine protease sequence against the custom database and identify its thermophilic nature.

📄 PDF Abstract BibTeX arXiv:2103.03512

Code (0)

등록된 구현이 없습니다.

Tasks

Classificationfeature selectionGeneral Classification

Similar Papers 제목 키워드 기반

AdaNovo: Adaptive \emph{De Novo} Peptide Sequencing with Conditional Mutual Information

2024-03-09 · Jun Xia, Shaorong Chen, Jingbo Zhou, Tianze Ling 외

Tandem mass spectrometry has played a pivotal role in advancing proteomics, enabling the analysis of protein composition in biological samples. Despite the development of various deep learning methods for identifying ami…

de novo peptide sequencing

NS-Pep: De novo Peptide Design with Non-Standard Amino Acids

2025-10-01 · Tao Guo, Junbo Yin, Yu Wang, Xin Gao arxiv

Peptide drugs incorporating non-standard amino acids (NSAAs) offer improved binding affinity and improved pharmacological properties. However, existing peptide design methods are limited to standard amino acids, leaving …

Xeno Amino Acids: A look into biochemistry as we don't know it

2023-10-24 · Sean M. Brown, Christopher Mayer-Bacon, Stephen Freeland

Would another origin of life resemble Earth's biochemical use of amino acids? Here we review current knowledge at three levels: 1) Could other classes of chemical structure serve as building blocks for biopolymer structu…

Classification of Macromolecule Type Based on Sequences of Amino Acids Using Deep Learning

2019-07-01 · Sarwar Khan, Faisal Ghaffar, Imad ali, qazi mazhar

The classification of amino acids and their sequence analysis plays a vital role in life sciences and is a challenging task. This article uses and compares state-of-the-art deep learning models like convolution neural ne…

General Classification

PepINVENT: Generative peptide design beyond the natural amino acids

2024-09-21 · Gökçe Geylan, Jon Paul Janet, Alessandro Tibo, Jiazhen He 외

Peptides play a crucial role in the drug design and discovery whether as a therapeutic modality or a delivery agent. Non-natural amino acids (NNAAs) have been used to enhance the peptide properties from binding affinity,…

Drug Design