paper-with-me

Papers

ProtCLIP: Function-Informed Protein Multi-Modal Learning

2024-12-28 · Hanjing Zhou, Mingze Yin, Wei Wu, Mingyang Li, Kun fu, Jintai Chen, Jian Wu, Zheng Wang

Multi-modality pre-training paradigm that aligns protein sequences and biological descriptions has learned general protein representations and achieved promising performance in various downstream applications. However, these works were still unable to replicate the extraordinary success of language-supervised visual foundation models due to the ineffective usage of aligned protein-text paired data and the lack of an effective function-informed pre-training paradigm. To address these issues, this paper curates a large-scale protein-text paired dataset called ProtAnno with a property-driven sampling strategy, and introduces a novel function-informed protein pre-training paradigm. Specifically, the sampling strategy determines selecting probability based on the sample confidence and property coverage, balancing the data quality and data quantity in face of large-scale noisy data. Furthermore, motivated by significance of the protein specific functional mechanism, the proposed paradigm explicitly model protein static and dynamic functional segments by two segment-wise pre-training objectives, injecting fine-grained information in a function-informed manner. Leveraging all these innovations, we develop ProtCLIP, a multi-modality foundation model that comprehensively represents function-aware protein embeddings. On 22 different protein benchmarks within 5 types, including protein functionality classification, mutation effect prediction, cross-modal transformation, semantic similarity inference and protein-protein interaction prediction, our ProtCLIP consistently achieves SOTA performance, with remarkable improvements of 75% on average in five cross-modal transformation benchmarks, 59.9% in GO-CC and 39.7% in GO-BP protein function prediction. The experimental results verify the extraordinary potential of ProtCLIP serving as the protein multi-modality foundation model.

📄 PDF Abstract BibTeX arXiv:2412.20014

Code (0)

등록된 구현이 없습니다.

Tasks

Protein Function PredictionSemantic SimilaritySemantic Textual Similarity

Similar Papers 제목 키워드 기반

Multi-Modal CLIP-Informed Protein Editing

2024-07-27 · Mingze Yin, Hanjing Zhou, Yiheng Zhu, Miao Lin 외

Proteins govern most biological functions essential for life, but achieving controllable protein discovery and optimization remains challenging. Recently, machine learning-assisted protein editing (MLPE) has shown promis…

AttributeContrastive Learning

Enhancing Protein-Protein Interaction Prediction with Hierarchical Motif-based Multimodal Protein Embedding

2026-05-30 · Zaifei Yang, Samuel Ping-Man Choi, James Kwok arxiv

Protein-protein interactions (PPIs) are essential for many biological processes. However, existing PPI prediction approaches suffer from two major limitations: they overlook the hierarchical organization of proteins, par…

MKGR: Multimodal Knowledge-Graph Representation Learning for Cold-Start Protein-Protein Interaction Prediction

2026-07-02 · Wenbo Zhang arxiv

Accurate protein-protein interaction (PPI) prediction is central to functional genomics, disease mechanism discovery, and drug development. A difficult setting arises when candidate interactions include proteins that hav…

Graph Representation LearningKnowledge GraphsGraph Learning

PRIME: Protein Representation via Physics-Informed Multiscale Equivariant Hierarchies

2026-05-02 · Viet Thanh Duy Nguyen, John K. Johnstone, Truong-Son Hy arxiv

Proteins are inherently multiscale physical systems whose functional properties emerge from coordinated structural organization across multiple spatial resolutions, ranging from atomic interactions to global fold topolog…

Representation Learning

Structure-Informed Protein Language Model

2024-02-07 · Zuobai Zhang, Jiarui Lu, Vijil Chenthamarakshan, Aurélie Lozano 외

Protein language models are a powerful tool for learning protein representations through pre-training on vast protein sequence datasets. However, traditional protein language models lack explicit structural supervision, …

Language ModelingLanguage ModellingmodelPrediction+2