Cross-Modality Protein Embedding for Compound-Protein Affinity and Contact Prediction
Compound-protein pairs dominate FDA-approved drug-target pairs and the prediction of compound-protein affinity and contact (CPAC) could help accelerate drug discovery. In this study we consider proteins as multi-modal data including 1D amino-acid sequences and (sequence-predicted) 2D residue-pair contact maps. We empirically evaluate the embeddings of the two single modalities in their accuracy and generalizability of CPAC prediction (i.e. structure-free interpretable compound-protein affinity prediction). And we rationalize their performances in both challenges of embedding individual modalities and learning generalizable embedding-label relationship. We further propose two models involving cross-modality protein embedding and establish that the one with cross interaction (thus capturing correlations among modalities) outperforms SOTAs and our single modality models in affinity, contact, and binding-site predictions for proteins never seen in the training set.
Code (0)
등록된 구현이 없습니다.
Tasks
Drug DiscoveryPredictionSimilar Papers 제목 키워드 기반
PSC-CPI: Multi-Scale Protein Sequence-Structure Contrasting for Efficient and Generalizable Compound-Protein Interaction Prediction
Compound-Protein Interaction (CPI) prediction aims to predict the pattern and strength of compound-protein interactions for rational drug discovery. Existing deep learning-based methods utilize only the single modality o…
Drug DiscoveryGraph neural networks and attention-based CNN-LSTM for protein classification
This paper focuses on three critical problems on protein classification. Firstly, Carbohydrate-active enzyme (CAZyme) classification can help people to understand the properties of enzymes. However, one CAZyme may belong…
ClassificationGraph AttentionGraph ClassificationGraph Learning+1TriFit: Trimodal Fusion with Protein Dynamics for Mutation Fitness Prediction
Predicting the functional impact of single amino acid substitutions (SAVs) is central to understanding genetic disease and engineering therapeutic proteins. While protein language models and structure-based methods have …
Contrastive LearningSeq2Mol: Automatic design of de novo molecules conditioned by the target protein sequences through deep neural networks
De novo design of molecules has recently enjoyed the power of generative deep neural networks. Current approaches aim to generate molecules either resembling the properties of the molecules of the training set or molecul…
Caption GenerationLanguage ModellingProtCLIP: Function-Informed Protein Multi-Modal Learning
Multi-modality pre-training paradigm that aligns protein sequences and biological descriptions has learned general protein representations and achieved promising performance in various downstream applications. However, t…
Protein Function PredictionSemantic SimilaritySemantic Textual Similarity