paper-with-me

Papers

MolTextNet: A Two-Million Molecule-Text Dataset for Multimodal Molecular Learning

2025-05-15 · Yihan Zhu, Gang Liu, Eric Inae, Meng Jiang

Small molecules are essential to drug discovery, and graph-language models hold promise for learning molecular properties and functions from text. However, existing molecule-text datasets are limited in scale and informativeness, restricting the training of generalizable multimodal models. We present MolTextNet, a dataset of 2.5 million high-quality molecule-text pairs designed to overcome these limitations. To construct it, we propose a synthetic text generation pipeline that integrates structural features, computed properties, bioactivity data, and synthetic complexity. Using GPT-4o-mini, we create structured descriptions for 2.5 million molecules from ChEMBL35, with text over 10 times longer than prior datasets. MolTextNet supports diverse downstream tasks, including property prediction and structure retrieval. Pretraining CLIP-style models with Graph Neural Networks and ModernBERT on MolTextNet yields improved performance, highlighting its potential for advancing foundational multimodal modeling in molecular science. Our dataset is available at https://huggingface.co/datasets/liuganghuggingface/moltextnet.

📄 PDF Abstract BibTeX arXiv:2506.00009

Code (0)

등록된 구현이 없습니다.

Tasks

Drug DiscoveryInformativenessProperty PredictionText Generation

Similar Papers 제목 키워드 기반

Molecular Identification from AFM images using the IUPAC Nomenclature and Attribute Multimodal Recurrent Neural Networks

2022-05-01 · Jaime Carracedo-Cosme, Carlos Romero-Muñiz, Pablo Pou, Rubén Pérez

Despite being the main tool to visualize molecules at the atomic scale, AFM with CO-functionalized metal tips is unable to chemically identify the observed molecules. Here we present a strategy to address this challengin…

AttributeImage Captioning

Otter-Knowledge: benchmarks of multimodal knowledge graph representation learning from different sources for drug discovery

2023-06-22 · Hoang Thanh Lam, Marco Luca Sbodio, Marcos Martínez Galindo, Mykhaylo Zayats 외

Recent research on predicting the binding affinity between drug molecules and proteins use representations learned, through unsupervised learning techniques, from large databases of molecule SMILES and protein sequences.…

Drug DiscoveryGraph Representation LearningKnowledge GraphsPrediction+1

MultiPUFFIN: A Multimodal Domain-Constrained Foundation Model for Molecular Property Prediction of Small Molecules

2026-03-01 · Idelfonso B. R. Nogueira, Carine M. Rebello, Mumin Enis Leblebici, Erick Giovani Sperandio Nascimento arxiv

MultiPUFFIN is a domain-informed multimodal foundation model for predicting thermophysical properties of small molecules, addressing a critical gap in chemical engineering, drug discovery, and materials science. Existing…

Molecular Property PredictionDrug Discovery

Hypothesis-and-Refinement Learning of Organic Structures from Multimodal Spectroscopic Data

2026-07-22 · Chengchun Liu, Zhiyuan Yan, Li Yuan, Hao Li 외 arxiv

Determining molecular structures from spectroscopic data remains fundamentally challenging because the inverse problem is intrinsically underdetermined: individual spectra are sparse, low-dimensional, and encode only par…

Automatic design of novel potential 3CL$^{\text{pro}}$ and PL$^{\text{pro}}$ inhibitors

2021-01-28 · Timothy Atkinson, Saeed Saremi, Faustino Gomez, Jonathan Masci

With the goal of designing novel inhibitors for SARS-CoV-1 and SARS-CoV-2, we propose the general molecule optimization framework, Molecular Neural Assay Search (MONAS), consisting of three components: a property predict…