paper-with-me

Papers

Does GNN Pretraining Help Molecular Representation?

2022-07-13 · Ruoxi Sun, Hanjun Dai, Adams Wei Yu

Extracting informative representations of molecules using Graph neural networks (GNNs) is crucial in AI-driven drug discovery. Recently, the graph research community has been trying to replicate the success of self-supervised pretraining in natural language processing, with several successes claimed. However, we find the benefit brought by self-supervised pretraining on small molecular data can be negligible in many cases. We conduct thorough ablation studies on the key components of GNN pretraining, including pretraining objectives, data splitting methods, input features, pretraining dataset scales, and GNN architectures, to see how they affect the accuracy of the downstream tasks. Our first important finding is, self-supervised graph pretraining do not always have statistically significant advantages over non-pretraining methods in many settings. Secondly, although noticeable improvement can be observed with additional supervised pretraining, the improvement may diminish with richer features or more balanced data splits. Thirdly, hyper-parameters could have larger impacts on accuracy of downstream tasks than the choice of pretraining tasks, especially when the scales of downstream tasks are small. Finally, we provide our conjectures where the complexity of some pretraining methods on small molecules might be insufficient, followed by empirical evidences on different pretraining datasets.

📄 PDF Abstract BibTeX arXiv:2207.06010

Code (0)

등록된 구현이 없습니다.

Tasks

Drug Discoverymolecular representation

Similar Papers 제목 키워드 기반

MolCAP: Molecular Chemical reActivity pretraining and prompted-finetuning enhanced molecular representation learning

2023-06-13 · Yu Wang, Jingjie Zhang, Junru Jin, Leyi Wei

Molecular representation learning (MRL) is a fundamental task for drug discovery. However, previous deep-learning (DL) methods focus excessively on learning robust inner-molecular representations by mask-dominated pretra…

DiversityDrug DiscoveryMolecular Property Predictionmolecular representation+2

Improving Molecular Pretraining with Complementary Featurizations

2022-09-29 · Yanqiao Zhu, Dingshuo Chen, Yuanqi Du, Yingze Wang 외

Molecular pretraining, which learns molecular representations over massive unlabeled data, has become a prominent paradigm to solve a variety of tasks in computational chemistry and drug discovery. Recently, prosperous p…

Computational chemistryDrug DiscoveryMolecular Property PredictionProperty Prediction

GeoRecon: Graph-Level Representation Learning for 3D Molecules via Reconstruction-Based Pretraining

2025-06-16 · Shaoheng Yan, Zian Li, Muhan Zhang

The pretraining-and-finetuning paradigm has driven significant advances across domains, such as natural language processing and computer vision, with representative pretraining paradigms such as masked language modeling …

DenoisingLanguage ModelingLanguage ModellingMasked Language Modeling+3

MoCL: Data-driven Molecular Fingerprint via Knowledge-aware Contrastive Learning from Molecular Graph

2021-06-05 · Mengying Sun, Jing Xing, Huijun Wang, Bin Chen 외

Recent years have seen a rapid growth of utilizing graph neural networks (GNNs) in the biomedical domain for tackling drug-related problems. However, like any other deep architectures, GNNs are data hungry. While requiri…

Contrastive LearningRepresentation Learning

3D Denoisers are Good 2D Teachers: Molecular Pretraining via Denoising and Cross-Modal Distillation

2023-09-08 · Sungjun Cho, Dae-Woong Jeong, Sung Moon Ko, Jinwoo Kim 외

Pretraining molecular representations from large unlabeled data is essential for molecular property prediction due to the high cost of obtaining ground-truth labels. While there exist various 2D graph-based molecular pre…

DenoisingKnowledge DistillationMolecular Property Predictionmolecular representation+2