paper-with-me

홈 › Papers

PepBenchmark: A Standardized Benchmark for Peptide Machine Learning

2026-04-12 · Jiahui Zhang, Rouyi Wang, Kuangqi Zhou, Tianshu Xiao, Lingyan Zhu, Yaosen Min, Yang Wang arxiv

Peptide therapeutics are widely regarded as the "third generation" of drugs, yet progress in peptide Machine Learning (ML) are hindered by the absence of standardized benchmarks. Here we present PepBenchmark, which unifies datasets, preprocessing, and evaluation protocols for peptide drug discovery. PepBenchmark comprises three components: (1) PepBenchData, a well-curated collection comprising 29 canonical-peptide and 6 non-canonical-peptide datasets across 7 groups, systematically covering key aspects of peptide drug development, representing, to the best of our knowledge, the most comprehensive AI-ready dataset resource to date; (2) PepBenchPipeline, a standardized preprocessing pipeline that ensures consistent dataset cleaning, construction, splitting, and feature transformation, mitigating quality issues common in ad hoc pipelines; and (3) PepBenchLeaderboard, a unified evaluation protocol and leaderboard with strong baselines across 4 major methodological families: Fingerprint-based, GNN-based, PLM-based, and SMILES-based models. Together, PepBenchmark provides the first standardized and comparable foundation for peptide drug discovery, facilitating methodological advances and translation into real-world applications. The data and code are publicly available at https://github.com/ZGCI-AI4S-Pep/PepBenchmark/.

📄 PDF Abstract BibTeX arXiv:2604.10531

Code (0)

등록된 구현이 없습니다.

Tasks

Drug Discovery

Similar Papers 제목 키워드 기반

A Standardized Benchmark for Multilabel Antimicrobial Peptide Classification

2025-11-06 · Sebastian Ojeda, Rafael Velasquez, Nicolás Aparicio, Juanita Puentes 외 arxiv

Antimicrobial peptides have emerged as promising molecules to combat antimicrobial resistance. However, fragmented datasets, inconsistent annotations, and the lack of standardized benchmarks hinder computational approach…

Accelerating MHC-II Epitope Discovery via Multi-Scale Prediction in Antigen Presentation

2025-12-16 · Yue Wan, Jiayi Yuan, Zhiwei Feng, Xiaowei Jia arxiv

Antigenic epitope presented by major histocompatibility complex II (MHC-II) proteins plays an essential role in immunotherapy. However, compared to the more widely studied MHC-I in computational immunotherapy, the study …

Topology-enhanced machine learning model (Top-ML) for anticancer peptide prediction

2024-07-12 · Joshua Zhi En Tan, JunJie Wee, Xue Gong, Kelin Xia

Recently, therapeutic peptides have demonstrated great promise for cancer treatment. To explore powerful anticancer peptides, artificial intelligence (AI)-based approaches have been developed to systematically screen pot…

Pep2Prob Benchmark: Predicting Fragment Ion Probability for MS$^2$-based Proteomics

2025-08-12 · Hao Xu, Zhichao Wang, Shengqi Sang, Pisit Wajanasara 외 arxiv

Proteins perform nearly all cellular functions and constitute most drug targets, making their analysis fundamental to understanding human biology in health and disease. Tandem mass spectrometry (MS$^2$) is the major anal…

Cross-Chirality Generalization by Axial Vectors for Hetero-Chiral Protein-Peptide Interaction Design

2026-02-13 · Ziyi Yang, Zitong Tian, Yinjun Jia, Tianyi Zhang 외 arxiv

D-peptide binders targeting L-proteins have promising therapeutic potential. Despite rapid advances in machine learning-based target-conditioned peptide design, generating D-peptide binders remains largely unexplored. In…

Protein Design