paper-with-me

홈 › Papers

MolPILE -- large-scale, diverse dataset for molecular representation learning

2025-09-22 · Jakub Adamczyk, Jakub Poziemski, Franciszek Job, Mateusz Król, Maciej Makowski arxiv

The size, diversity, and quality of pretraining datasets critically determine the generalization ability of foundation models. Despite their growing importance in chemoinformatics, the effectiveness of molecular representation learning has been hindered by limitations in existing small molecule datasets. To address this gap, we present MolPILE, large-scale, diverse, and rigorously curated collection of 222 million compounds, constructed from 6 large-scale databases using an automated curation pipeline. We present a comprehensive analysis of current pretraining datasets, highlighting considerable shortcomings for training ML models, and demonstrate how retraining existing models on MolPILE yields improvements in generalization performance. This work provides a standardized resource for model training, addressing the pressing need for an ImageNet-like dataset in molecular chemistry.

📄 PDF Abstract BibTeX arXiv:2509.18353

Code (0)

등록된 구현이 없습니다.

Tasks

Representation Learning

Similar Papers 제목 키워드 기반

Large-Scale Knowledge Integration for Enhanced Molecular Property Prediction

2024-10-15 · Yasir Ghunaim, Robert Hoehndorf

Pre-training machine learning models on molecular properties has proven effective for generating robust and generalizable representations, which is critical for advancements in drug discovery and materials science. While…

Drug DiscoveryMolecular Property PredictionPredictionProperty Prediction

Mol-LLM: Multimodal Generalist Molecular LLM with Improved Graph Utilization

2025-02-05 · Chanhui Lee, Hanbum Ko, Yuheon Song, Yongjun Jeong 외

Recent advances in large language models (LLMs) have led to models that tackle diverse molecular tasks, such as chemical reaction prediction and molecular property prediction. Large-scale molecular instruction-tuning dat…

Chemical Reaction PredictionMolecular Property PredictionMolecule CaptioningPrediction+1

Smiles2Dock: an open large-scale multi-task dataset for ML-based molecular docking

2024-06-09 · Thomas Le Menestrel, Manuel Rivas

Docking is a crucial component in drug discovery aimed at predicting the binding conformation and affinity between small molecules and target proteins. ML-based docking has recently emerged as a prominent approach, outpa…

BenchmarkingDrug DiscoveryMolecular Docking

M$^{3}$-20M: A Large-Scale Multi-Modal Molecule Dataset for AI-driven Drug Design and Discovery

2024-12-08 · Siyuan Guo, Lexuan Wang, Chang Jin, Jinxian Wang 외

This paper introduces M$^{3}$-20M, a large-scale Multi-Modal Molecule dataset that contains over 20 million molecules, with the data mainly being integrated from existing databases and partially generated by large langua…

Drug DesignMolecular Property PredictionProperty Predictionvalid

DG-GL: Differential geometry based geometric learning of molecular datasets

2018-06-11

Motivation: Despite its great success in various physical modeling, differential geometry (DG) has rarely been devised as a versatile tool for analyzing large, diverse and complex molecular and biomolecular datasets due …

DescriptiveDimensionality ReductionDrug Discovery