paper-with-me

홈 › Papers

NaFM: Pre-training a Foundation Model for Small-Molecule Natural Products

2025-03-22 · Yuheng Ding, Bo Qiang, Yiran Zhou, Jie Yu, Qi Li, Liangren Zhang, Yusong Wang, Zhenmin Liu

Natural products, as metabolites from microorganisms, animals, or plants, exhibit diverse biological activities, making them crucial for drug discovery. Nowadays, existing deep learning methods for natural products research primarily rely on supervised learning approaches designed for specific downstream tasks. However, such one-model-for-a-task paradigm often lacks generalizability and leaves significant room for performance improvement. Additionally, existing molecular characterization methods are not well-suited for the unique tasks associated with natural products. To address these limitations, we have pre-trained a foundation model for natural products based on their unique properties. Our approach employs a novel pretraining strategy that is especially tailored to natural products. By incorporating contrastive learning and masked graph learning objectives, we emphasize evolutional information from molecular scaffolds while capturing side-chain information. Our framework achieves state-of-the-art (SOTA) results in various downstream tasks related to natural product mining and drug discovery. We first compare taxonomy classification with synthesized molecule-focused baselines to demonstrate that current models are inadequate for understanding natural synthesis. Furthermore, by diving into a fine-grained analysis at both the gene and microbial levels, NaFM demonstrates the ability to capture evolutionary information. Eventually, our method is experimented with virtual screening, illustrating informative natural product representations that can lead to more effective identification of potential drug candidates.

📄 PDF Abstract BibTeX arXiv:2503.17656

Code (1)

tomaidd/nafm-official 공식 구현 pytorch

Tasks

Contrastive LearningDrug DiscoveryGraph Learning

Methods 이 논문이 사용한 방법론

Contrastive Learning 설명 없음

Similar Papers 제목 키워드 기반

FARM: Functional Group-Aware Representations for Small Molecules

2024-10-02 · Thao Nguyen, Kuan-Hao Huang, Ge Liu, Martin D. Burke 외

We introduce Functional Group-Aware Representations for Small Molecules (FARM), a novel foundation model designed to bridge the gap between SMILES, natural language, and molecular graphs. The key innovation of FARM lies …

Contrastive LearningDrug DiscoveryLanguage ModelingLanguage Modelling+3

MultiPUFFIN: A Multimodal Domain-Constrained Foundation Model for Molecular Property Prediction of Small Molecules

2026-03-01 · Idelfonso B. R. Nogueira, Carine M. Rebello, Mumin Enis Leblebici, Erick Giovani Sperandio Nascimento arxiv

MultiPUFFIN is a domain-informed multimodal foundation model for predicting thermophysical properties of small molecules, addressing a critical gap in chemical engineering, drug discovery, and materials science. Existing…

Molecular Property PredictionDrug Discovery

A Systematic Evaluation of Co-folding Model Representations for Small-Molecule Learning

2026-02-02 · Hyosoon Jang, Hyunjin Seo, Honghui Kim, Seonghyun Park 외 arxiv

Small-molecule foundation models are typically pretrained on standalone molecular data, unlike vision and language models that often benefit from cross-modal or relational supervision. Protein-ligand co-folding provides …

Representation LearningReinforcement Learning

NatureLM: Deciphering the Language of Nature for Scientific Discovery

2025-02-11 · Yingce Xia, Peiran Jin, Shufang Xie, Liang He 외

Foundation models have revolutionized natural language processing and artificial intelligence, significantly enhancing how machines comprehend and generate human languages. Inspired by the success of these foundation mod…

Drug DiscoveryRetrosynthesisscientific discovery

L+M-24: Building a Dataset for Language + Molecules @ ACL 2024

2024-02-22 · Carl Edwards, Qingyun Wang, Lawrence Zhao, Heng Ji

Language-molecule models have emerged as an exciting direction for molecular discovery and understanding. However, training these models is challenging due to the scarcity of molecule-language pair datasets. At this poin…

Entity LinkingProperty Prediction