paper-with-me

홈 › Papers

Structure-Enhanced Protein Instruction Tuning: Towards General-Purpose Protein Understanding with LLMs

2024-10-04 · Wei Wu, Chao Wang, Liyi Chen, Mingze Yin, Yiheng Zhu, Kun fu, Jieping Ye, Hui Xiong, Zheng Wang

Proteins, as essential biomolecules, play a central role in biological processes, including metabolic reactions and DNA replication. Accurate prediction of their properties and functions is crucial in biological applications. Recent development of protein language models (pLMs) with supervised fine tuning provides a promising solution to this problem. However, the fine-tuned model is tailored for particular downstream prediction task, and achieving general-purpose protein understanding remains a challenge. In this paper, we introduce Structure-Enhanced Protein Instruction Tuning (SEPIT) framework to bridge this gap. Our approach incorporates a novel structure-aware module into pLMs to enrich their structural knowledge, and subsequently integrates these enhanced pLMs with large language models (LLMs) to advance protein understanding. In this framework, we propose a novel instruction tuning pipeline. First, we warm up the enhanced pLMs using contrastive learning and structure denoising. Then, caption-based instructions are used to establish a basic understanding of proteins. Finally, we refine this understanding by employing a mixture of experts (MoEs) to capture more complex properties and functional information with the same number of activated parameters. Moreover, we construct the largest and most comprehensive protein instruction dataset to date, which allows us to train and evaluate the general-purpose protein understanding model. Extensive experiments on both open-ended generation and closed-set answer tasks demonstrate the superior performance of SEPIT over both closed-source general LLMs and open-source LLMs trained with protein knowledge.

📄 PDF Abstract BibTeX arXiv:2410.03553

Code (0)

등록된 구현이 없습니다.

Tasks

Contrastive LearningDenoisingMixture-of-Experts

Methods 이 논문이 사용한 방법론

Contrastive Learning 설명 없음

Similar Papers 제목 키워드 기반

EvoLlama: Enhancing LLMs' Understanding of Proteins via Multimodal Structure and Sequence Representations

2024-12-16 · Nuowei Liu, Changzhi Sun, Tao Ji, Junfeng Tian 외

Current Large Language Models (LLMs) for understanding proteins primarily treats amino acid sequences as a text modality. Meanwhile, Protein Language Models (PLMs), such as ESM-2, have learned massive sequential evolutio…

Property Prediction

ProLLM: Protein Chain-of-Thoughts Enhanced LLM for Protein-Protein Interaction Prediction

2024-03-30 · Mingyu Jin, Haochen Xue, Zhenting Wang, Boming Kang 외

The prediction of protein-protein interactions (PPIs) is crucial for understanding biological functions and diseases. Previous machine learning approaches to PPI prediction mainly focus on direct physical interactions, i…

ProteinGPT: Multimodal LLM for Protein Property Prediction and Structure Understanding

2024-08-21 · Yijia Xiao, Edward Sun, Yiqiao Jin, Qifan Wang 외

Understanding biological processes, drug development, and biotechnological advancements requires a detailed analysis of protein structures and functions, a task that is inherently complex and time-consuming in traditiona…

Language ModelingLanguage ModellingLarge Language ModelMultimodal Large Language Model+1

STELLA: Towards Protein Function Prediction with Multimodal LLMs Integrating Sequence-Structure Representations

2025-06-04 · Hongwang Xiao, Wenjun Lin, Xi Chen, Hui Wang 외

Protein biology focuses on the intricate relationships among sequences, structures, and functions. Deciphering protein functions is crucial for understanding biological processes, advancing drug discovery, and enabling s…

Drug DiscoveryGeneral KnowledgePredictionProtein Function Prediction

TourSynbio: A Multi-Modal Large Model and Agent Framework to Bridge Text and Protein Sequences for Protein Engineering

2024-08-27 · Yiqing Shen, Zan Chen, Michail Mamalakis, Yungeng Liu 외

The structural similarities between protein sequences and natural languages have led to parallel advancements in deep learning across both domains. While large language models (LLMs) have achieved much progress in the do…

Multiple-choiceProtein Folding