paper-with-me

홈 › Papers

Prot2Chat: Protein LLM with Early-Fusion of Text, Sequence and Structure

2025-02-07 · Zhicong Wang, Zicheng Ma, Ziqiang Cao, Changlong Zhou, Jun Zhang, Yiqin Gao

Motivation: Proteins are of great significance in living organisms. However, understanding their functions encounters numerous challenges, such as insufficient integration of multimodal information, a large number of training parameters, limited flexibility of classification-based methods, and the lack of systematic evaluation metrics for protein Q&A systems. To tackle these issues, we propose the Prot2Chat framework. Results: We modified ProteinMPNN to encode protein sequence and structural information in a unified way. We used a large language model (LLM) to encode questions into vectors and developed a protein-text adapter to compress protein information into virtual tokens based on these vectors, achieving the early fusion of text and protein information. Finally, the same LLM reads the virtual tokens and the questions to generate answers. To optimize training efficiency, we froze the encoder and employed Low-Rank Adaptation (LoRA) techniques for the LLM. Experiments on two datasets show that both automated metrics and expert evaluations demonstrate the superior performance of our model, and zero-shot prediction results highlight its generalization ability. The models and codes are available at https://github.com/ wangzc1233/Prot2Chat. Contact: zqcao@suda.edu.cn or wangzc025@163.com Key words: Protein Q&A, Early-Fusion, LLM

📄 PDF Abstract BibTeX arXiv:2502.06846

Code (1)

wangzc1233/Prot2Chat 공식 구현 pytorch

Tasks

Answer GenerationDecoderLanguage ModelingLanguage ModellingLarge Language ModelNatural Language Understanding

Methods 이 논문이 사용한 방법론

Adapter 설명 없음

Similar Papers 제목 키워드 기반

ProtChatGPT: Towards Understanding Proteins with Large Language Models

2024-02-15 · Chao Wang, Hehe Fan, Ruijie Quan, Yi Yang

Protein research is crucial in various fundamental disciplines, but understanding their intricate structure-function relationships remains challenging. Recent Large Language Models (LLMs) have made significant strides in…

GEFA: Early Fusion Approach in Drug-Target Affinity Prediction

2020-09-25 · Tri Minh Nguyen, Thin Nguyen, Thao Minh Le, Truyen Tran

Predicting the interaction between a compound and a target is crucial for rapid drug repurposing. Deep learning has been successfully applied in drug-target affinity (DTA) problem. However, previous deep learning-based m…

Graph Neural Network

Steering Masked Discrete Diffusion Models via Discrete Denoising Posterior Prediction

2024-10-10 · Jarrid Rector-Brooks, Mohsin Hasan, Zhangzhi Peng, Zachary Quinn 외

Generative modeling of discrete data underlies important applications spanning text-based agents like ChatGPT to the design of the very building blocks of life in protein sequences. However, application domains need to e…

Denoising

Ligand-Conditioned Discrete Diffusion for Protein Sequence-Structure Co-Design

2026-05-15 · Chen Wei, Fanding Xu, Minghao Sun, Zhiyuan Liu 외 arxiv

Proteins perform their biological functions through three-dimensional structures encoded by amino acid sequences, and ligand-binding protein co-design requires models that generate sequence-structure compatible proteins …

Protein Design

Protein generation with embedding learning for motif diversification

2025-10-21 · Kevin Michalewicz, Chen Jin, Philip Alexander Teare, Tom Diethe 외 arxiv

A fundamental challenge in protein design is the trade-off between generating structural diversity while preserving motif biological function. Current state-of-the-art methods, such as partial diffusion in RFdiffusion, o…

Protein Design