paper-with-me

홈 › Papers

A Large Encoder-Decoder Family of Foundation Models For Chemical Language

2024-07-24 · Eduardo Soares, Victor Shirasuna, Emilio Vital Brazil, Renato Cerqueira, Dmitry Zubarev, Kristin Schmidt

Large-scale pre-training methodologies for chemical language models represent a breakthrough in cheminformatics. These methods excel in tasks such as property prediction and molecule generation by learning contextualized representations of input tokens through self-supervised learning on large unlabeled corpora. Typically, this involves pre-training on unlabeled data followed by fine-tuning on specific tasks, reducing dependence on annotated datasets and broadening chemical language representation understanding. This paper introduces a large encoder-decoder chemical foundation models pre-trained on a curated dataset of 91 million SMILES samples sourced from PubChem, which is equivalent to 4 billion of molecular tokens. The proposed foundation model supports different complex tasks, including quantum property prediction, and offer flexibility with two main variants (289M and $8\times289M$). Our experiments across multiple benchmark datasets validate the capacity of the proposed model in providing state-of-the-art results for different tasks. We also provide a preliminary assessment of the compositionality of the embedding space as a prerequisite for the reasoning tasks. We demonstrate that the produced latent space is separable compared to the state-of-the-art with few-shot learning capabilities.

📄 PDF Abstract BibTeX arXiv:2407.20267

Code (1)

🤗 ibm/materials.smi-ted 공식 구현

Tasks

DecoderFew-Shot LearningProperty PredictionSelf-Supervised Learning

Similar Papers 제목 키워드 기반

Improving Chemical Autoencoder Latent Space and Molecular De novo Generation Diversity with Heteroencoders

2018-06-25 · Esben Jannik Bjerrum, Boris Sattarov

Chemical autoencoders are attractive models as they combine chemical space navigation with possibilities for de-novo molecule generation in areas of interest. This enables them to produce focused chemical libraries aroun…

DecoderDiversityDrug Discovery

From Syntax to Semantics: Unveiling the Emergence of Chirality in SMILES Translation Models

2026-05-11 · Zehao Li, Yasuhiro Yoshikai, Shumpei Nemoto, Hiroyuki Kusuhara 외 arxiv

Understanding how chemical language models (CLMs) learn chemical meaning from molecular string representations, rather than only surface-level string patterns, is an important question in chemical representation learning…

Representation Learning

nach0: Multimodal Natural and Chemical Languages Foundation Model

2023-11-21 · Micha Livne, Zulfat Miftahutdinov, Elena Tutubalina, Maksim Kuznetsov 외

Large Language Models (LLMs) have substantially driven scientific progress in various domains, and many papers have demonstrated their ability to tackle complex problems with creative solutions. Our paper introduces a ne…

Decodermodelnamed-entity-recognitionNamed Entity Recognition+1

Geometric Foundation Model Distillation for Efficient Lunar 3D Reconstruction

2026-07-02 · Clémentine Grethen, Florient Chouteau, Géraldine Morin, Simone Gasparini arxiv

Large 3D foundation models such as MASt3R achieve state-of-the-art stereo reconstruction but are computationally demanding for deployment under strict hardware constraints -- a critical limitation in domains such as plan…

Knowledge Distillation3D Reconstruction

Foundation Models for Discovery and Exploration in Chemical Space

2025-10-20 · Alexius Wadell, Anoushka Bhutani, Victor Azumah, Austin R. Ellis-Mohr 외 arxiv

Accurate prediction of atomistic, thermodynamic, and kinetic properties from molecular structures underpins materials innovation. Existing computational and experimental approaches lack the scalability required to naviga…