paper-with-me

홈 › Papers

Multi-modal Molecule Structure-text Model for Text-based Retrieval and Editing

2022-12-21 · Shengchao Liu, Weili Nie, Chengpeng Wang, Jiarui Lu, Zhuoran Qiao, Ling Liu, Jian Tang, Chaowei Xiao, Anima Anandkumar

There is increasing adoption of artificial intelligence in drug discovery. However, existing studies use machine learning to mainly utilize the chemical structures of molecules but ignore the vast textual knowledge available in chemistry. Incorporating textual knowledge enables us to realize new drug design objectives, adapt to text-based instructions and predict complex biological activities. Here we present a multi-modal molecule structure-text model, MoleculeSTM, by jointly learning molecules' chemical structures and textual descriptions via a contrastive learning strategy. To train MoleculeSTM, we construct a large multi-modal dataset, namely, PubChemSTM, with over 280,000 chemical structure-text pairs. To demonstrate the effectiveness and utility of MoleculeSTM, we design two challenging zero-shot tasks based on text instructions, including structure-text retrieval and molecule editing. MoleculeSTM has two main properties: open vocabulary and compositionality via natural language. In experiments, MoleculeSTM obtains the state-of-the-art generalization ability to novel biochemical concepts across various benchmarks.

📄 PDF Abstract BibTeX arXiv:2212.10789

Code (1)

chao1224/moleculestm 공식 구현 pytorch

Tasks

Contrastive LearningDrug DesignDrug DiscoveryRetrievalText Retrieval

Methods 이 논문이 사용한 방법론

Contrastive Learning 설명 없음

Similar Papers 제목 키워드 기반

MolTextNet: A Two-Million Molecule-Text Dataset for Multimodal Molecular Learning

2025-05-15 · Yihan Zhu, Gang Liu, Eric Inae, Meng Jiang

Small molecules are essential to drug discovery, and graph-language models hold promise for learning molecular properties and functions from text. However, existing molecule-text datasets are limited in scale and informa…

Drug DiscoveryInformativenessProperty PredictionText Generation

GeomCLIP: Contrastive Geometry-Text Pre-training for Molecules

2024-11-16 · Teng Xiao, Chao Cui, Huaisheng Zhu, Vasant G. Honavar

Pretraining molecular representations is crucial for drug and material discovery. Recent methods focus on learning representations from geometric structures, effectively capturing 3D position information. Yet, they overl…

DenoisingMolecular Property PredictionMolecule CaptioningProperty Prediction+1

Adversarial Modality Alignment Network for Cross-Modal Molecule Retrieval

2023-03-08 · IEEE Transactions on Artificial Intelligence 2023 3 · Wenyu Zhao, Dong Zhou, Buqing Cao, Kai Zhang 외

The cross-modal molecule retrieval (Text2Mol) task aims to bridge the semantic gap between molecules and natural language descriptions. A solution to this non-trivial problem relies on graph convolutional network (GCN) a…

Contrastive LearningCross-Modal RetrievalRetrievalTriplet

Exploring Optimal Transport-Based Multi-Grained Alignments for Text-Molecule Retrieval

2024-11-04 · Zijun Min, Bingshuai Liu, Liang Zhang, Jia Song 외

The field of bioinformatics has seen significant progress, making the cross-modal text-molecule retrieval task increasingly vital. This task focuses on accurately retrieving molecule structures based on textual descripti…

Contrastive LearningCross-Modal RetrievalRetrievalSentence

Text2Mol: Cross-Modal Molecule Retrieval with Natural Language Queries

2021-11-01 · EMNLP 2021 11 · Carl Edwards, ChengXiang Zhai, Heng Ji

We propose a new task, Text2Mol, to retrieve molecules using natural language descriptions as queries. Natural language and molecules encode information in very different ways, which leads to the exciting but challenging…

Cross-Modal RetrievalNatural Language QueriesRerankingRetrieval