Deep Learning Model for Amyloidogenicity Prediction using a Pre-trained Protein LLM
The prediction of amyloidogenicity in peptides and proteins remains a focal point of ongoing bioinformatics. The crucial step in this field is to apply advanced computational methodologies. Many recent approaches to predicting amyloidogenicity within proteins are highly based on evolutionary motifs and the individual properties of amino acids. It is becoming increasingly evident that the sequence information-based features show high predictive performance. Consequently, our study evaluated the contextual features of protein sequences obtained from a pretrained protein large language model leveraging bidirectional LSTM and GRU to predict amyloidogenic regions in peptide and protein sequences. Our method achieved an accuracy of 84.5% on 10-fold cross-validation and an accuracy of 83% in the test dataset. Our results demonstrate competitive performance, highlighting the potential of LLMs in enhancing the accuracy of amyloid prediction.
Code (0)
등록된 구현이 없습니다.
Similar Papers 제목 키워드 기반
PEER: A Comprehensive and Multi-Task Benchmark for Protein Sequence Understanding
We are now witnessing significant progress of deep learning methods in a variety of tasks (or datasets) of proteins. However, there is a lack of a standard benchmark to evaluate the performance of different methods, whic…
Feature EngineeringMulti-Task LearningPredictionProtein Function Prediction+1Random Embeddings and Linear Regression can Predict Protein Function
Large self-supervised models pretrained on millions of protein sequences have recently gained popularity in generating embeddings of protein sequences for protein function prediction. However, the absence of random basel…
PredictionProtein Function PredictionregressionExploring Large Protein Language Models in Constrained Evaluation Scenarios within the FLIP Benchmark
In this study, we expand upon the FLIP benchmark-designed for evaluating protein fitness prediction models in small, specialized prediction tasks-by assessing the performance of state-of-the-art large protein language mo…
PredictionIntegration of persistent Laplacian and pre-trained transformer for protein solubility changes upon mutation
Protein mutations can significantly influence protein solubility, which results in altered protein functions and leads to various diseases. Despite of tremendous effort, machine learning prediction of protein solubility …
Integration of Pre-trained Protein Language Models into Geometric Deep Learning Networks
Geometric deep learning has recently achieved great success in non-Euclidean domains, and learning on 3D structures of large biomolecules is emerging as a distinct research area. However, its efficacy is largely constrai…
Protein Interface PredictionRepresentation Learning