Robust Molecular Property Prediction via Densifying Scarce Labeled Data
A widely recognized limitation of molecular prediction models is their reliance on structures observed in the training data, resulting in poor generalization to out-of-distribution compounds. Yet in drug discovery, the compounds most critical for advancing research often lie beyond the training set, making the bias toward the training data particularly problematic. This mismatch introduces substantial covariate shift, under which standard deep learning models produce unstable and inaccurate predictions. Furthermore, the scarcity of labeled data, stemming from the onerous and costly nature of experimental validation, further exacerbates the difficulty of achieving reliable generalization. To address these limitations, we propose a novel meta-learning-based approach that leverages unlabeled data to interpolate between in-distribution (ID) and out-of-distribution (OOD) data, enabling the model to meta-learn how to generalize beyond the training distribution. We demonstrate significant performance gains over state-of-the-art methods on challenging real-world datasets that exhibit substantial covariate shift.
Code (1)
Tasks
Drug DiscoveryMeta-LearningMolecular Property PredictionProperty PredictionSimilar Papers 제목 키워드 기반
ASGN: An Active Semi-supervised Graph Neural Network for Molecular Property Prediction
Molecular property prediction (e.g., energy) is an essential problem in chemistry and biology. Unfortunately, many supervised learning methods usually suffer from the problem of scarce labeled molecules in the chemical s…
Active LearningGraph Neural NetworkMolecular Property PredictionProperty PredictionTwo-Stage Pretraining for Molecular Property Prediction in the Wild
Accurate property prediction is crucial for accelerating the discovery of new molecules. Although deep learning models have achieved remarkable success, their performance often relies on large amounts of labeled data tha…
DenoisingMolecular Property PredictionProperty PredictionCardinality-Preserving Attention Channels for Graph Transformers in Molecular Property Prediction
Molecular property prediction is crucial for drug discovery when labeled data are scarce. This work presents CardinalGraphFormer, a graph transformer augmented with a query-conditioned cardinality-preserving attention (C…
Molecular Property PredictionDrug DiscoveryPCEvo: Path-Consistent Molecular Representation via Virtual Evolutionary
Molecular representation learning aims to learn vector embeddings that capture molecular structure and geometry, thereby enabling property prediction and downstream scientific applications. In many AI for science tasks, …
Representation LearningImproving VAE based molecular representations for compound property prediction
Collecting labeled data for many important tasks in chemoinformatics is time consuming and requires expensive experiments. In recent years, machine learning has been used to learn rich representations of molecules using …
BIG-bench Machine LearningPredictionProperty Prediction