paper-with-me

홈 › Papers

ChemPile: A 250GB Diverse and Curated Dataset for Chemical Foundation Models

2025-05-18 · Adrian Mirza, Nawaf Alampara, Martiño Ríos-García, Mohamed Abdelalim, Jack Butler, Bethany Connolly, Tunca Dogan, Marianna Nezhurina, Bünyamin Şen, Santosh Tirunagari, Mark Worrall, Adamo Young, Philippe Schwaller, Michael Pieler, Kevin Maik Jablonka

Foundation models have shown remarkable success across scientific domains, yet their impact in chemistry remains limited due to the absence of diverse, large-scale, high-quality datasets that reflect the field's multifaceted nature. We present the ChemPile, an open dataset containing over 75 billion tokens of curated chemical data, specifically built for training and evaluating general-purpose models in the chemical sciences. The dataset mirrors the human learning journey through chemistry -- from educational foundations to specialized expertise -- spanning multiple modalities and content types including structured data in diverse chemical representations (SMILES, SELFIES, IUPAC names, InChI, molecular renderings), scientific and educational text, executable code, and chemical images. ChemPile integrates foundational knowledge (textbooks, lecture notes), specialized expertise (scientific articles and language-interfaced data), visual understanding (molecular structures, diagrams), and advanced reasoning (problem-solving traces and code) -- mirroring how human chemists develop expertise through diverse learning materials and experiences. Constructed through hundreds of hours of expert curation, the ChemPile captures both foundational concepts and domain-specific complexity. We provide standardized training, validation, and test splits, enabling robust benchmarking. ChemPile is openly released via HuggingFace with a consistent API, permissive license, and detailed documentation. We hope the ChemPile will serve as a catalyst for chemical AI, enabling the development of the next generation of chemical foundation models.

📄 PDF Abstract BibTeX arXiv:2505.12534

Code (0)

등록된 구현이 없습니다.

Tasks

ArticlesBenchmarking

Similar Papers 제목 키워드 기반

Bi-semantic Chemical Embedder for Joint Representation Learning of SMILES and Natural Language

2026-08-04 · David Ming Segura, Jeremy Goumaz, Joshua W. Sin, Bojana Ranković 외 arxiv

Transformer models have revolutionized natural language processing (NLP), and text-based molecular representations like SMILES have successfully extended these architectures to chemistry. However, domain-adaptive pre-tra…

Molecular Property PredictionRepresentation LearningContrastive Learning

A Large Encoder-Decoder Family of Foundation Models For Chemical Language

2024-07-24 · Eduardo Soares, Victor Shirasuna, Emilio Vital Brazil, Renato Cerqueira 외

Large-scale pre-training methodologies for chemical language models represent a breakthrough in cheminformatics. These methods excel in tasks such as property prediction and molecule generation by learning contextualized…

DecoderFew-Shot LearningProperty PredictionSelf-Supervised Learning

Generative Chemical Language Models for Energetic Materials Discovery

2026-03-30 · Andrew Salij, R. Seaton Ullberg, Megan C. Davis, Marc J. Cawkwell 외 arxiv

The discovery of new energetic materials remains a pressing challenge hindered by limited availability of high-quality data. To address this, we have developed generative molecular language models that have been pretrain…

HELM-BERT: A Transformer for Medium-sized Peptide Property Prediction

2025-12-29 · Seungeon Lee, Takuto Koyama, Itsuki Maeda, Shigeyuki Matsumoto 외 arxiv

Therapeutic peptides have emerged as a pivotal modality in modern drug discovery, occupying a chemically and topologically rich space. While accurate prediction of their physicochemical properties is essential for accele…

Drug Discovery

GlycoNMR: Dataset and benchmarks for NMR chemical shift prediction of carbohydrates with graph neural networks

2023-11-28 · Zizhang Chen, Ryan Paul Badman, Lachele Foley, Robert Woods 외

Molecular representation learning (MRL) is a powerful tool for bridging the gap between machine learning and chemical sciences, as it converts molecules into numerical representations while preserving their chemical feat…

Drug Designmolecular representationProperty PredictionRepresentation Learning