paper-with-me

홈 › Papers

PepDoRA: A Unified Peptide Language Model via Weight-Decomposed Low-Rank Adaptation

2024-10-28 · Leyao Wang, Rishab Pulugurta, Pranay Vure, Yinuo Zhang, Aastha Pal, Pranam Chatterjee

Peptide therapeutics, including macrocycles, peptide inhibitors, and bioactive linear peptides, play a crucial role in therapeutic development due to their unique physicochemical properties. However, predicting these properties remains challenging. While structure-based models primarily focus on local interactions, language models are capable of capturing global therapeutic properties of both modified and linear peptides. Protein language models like ESM-2, though effective for natural peptides, cannot however encode chemical modifications. Conversely, pre-trained chemical language models excel in representing small molecule properties but are not optimized for peptides. To bridge this gap, we introduce PepDoRA, a unified peptide representation model. Leveraging Weight-Decomposed Low-Rank Adaptation (DoRA), PepDoRA efficiently fine-tunes the ChemBERTa-77M-MLM on a masked language model objective to generate optimized embeddings for downstream property prediction tasks involving both modified and unmodified peptides. By tuning on a diverse and experimentally valid set of 100,000 modified, bioactive, and binding peptides, we show that PepDoRA embeddings capture functional properties of input peptides, enabling the accurate prediction of membrane permeability, non-fouling and hemolysis propensity, and via contrastive learning, target protein-specific binding. Overall, by providing a unified representation for chemically and biologically diverse peptides, PepDoRA serves as a versatile tool for function and activity prediction, facilitating the development of peptide therapeutics across a broad spectrum of applications.

📄 PDF Abstract BibTeX arXiv:2410.20667

Code (0)

등록된 구현이 없습니다.

Tasks

Activity PredictionContrastive LearningLanguage ModelingLanguage ModellingProperty Prediction

Methods 이 논문이 사용한 방법론

SET Dynamic Sparse Training method where weight mask is updated randomly periodically
Focus 설명 없음

Similar Papers 제목 키워드 기반

HELM-BERT: A Transformer for Medium-sized Peptide Property Prediction

2025-12-29 · Seungeon Lee, Takuto Koyama, Itsuki Maeda, Shigeyuki Matsumoto 외 arxiv

Therapeutic peptides have emerged as a pivotal modality in modern drug discovery, occupying a chemically and topologically rich space. While accurate prediction of their physicochemical properties is essential for accele…

Drug Discovery

Training self-supervised peptide sequence models on artificially chopped proteins

2022-11-09 · Gil Sadeh, Zichen Wang, Jasleen Grewal, Huzefa Rangwala 외

Representation learning for proteins has primarily focused on the global understanding of protein sequences regardless of their length. However, shorter proteins (known as peptides) take on distinct structures and functi…

Data AugmentationLanguage ModelingLanguage ModellingRepresentation Learning+1

NS-Pep: De novo Peptide Design with Non-Standard Amino Acids

2025-10-01 · Tao Guo, Junbo Yin, Yu Wang, Xin Gao arxiv

Peptide drugs incorporating non-standard amino acids (NSAAs) offer improved binding affinity and improved pharmacological properties. However, existing peptide design methods are limited to standard amino acids, leaving …

DapPep: Domain Adaptive Peptide-agnostic Learning for Universal T-cell Receptor-antigen Binding Affinity Prediction

2024-11-26 · Jiangbin Zheng, Qianhui Xu, Ruichen Xia, Stan Z. Li

Identifying T-cell receptors (TCRs) that interact with antigenic peptides provides the technical basis for developing vaccines and immunotherapies. The emergent deep learning methods excel at learning antigen binding pat…

Language ModelingLanguage ModellingProtein Language ModelSpecificity

PepBenchmark: A Standardized Benchmark for Peptide Machine Learning

2026-04-12 · Jiahui Zhang, Rouyi Wang, Kuangqi Zhou, Tianshu Xiao 외 arxiv

Peptide therapeutics are widely regarded as the "third generation" of drugs, yet progress in peptide Machine Learning (ML) are hindered by the absence of standardized benchmarks. Here we present PepBenchmark, which unifi…

Drug Discovery