paper-with-me

Papers

Molecular Representations for Large Language Models

2026-05-03 · Nicholas T. Runcie, Fergus Imrie, Charlotte M. Deane arxiv

Large Language Models (LLMs) are increasingly being used to support scientific discovery. In chemistry, tasks such as reaction prediction and structure elucidation require reasoning about the structures of molecules. As such, LLM-based systems for chemistry must interact reliably with molecular structures. Most previous studies of LLMs in chemistry have used SMILES strings or IUPAC names as molecular representations; however, the suitability of these formats has not been systematically assessed. In this work, we introduce MolJSON, a novel molecular representation for LLMs, and systematically compare it with five common chemical formats. We evaluated each representation with GPT-5-nano, GPT-5-mini, GPT-5, and Claude Haiku 4.5 using a set of 78,045 questions spanning translation, shortest path, and constrained generation reasoning tasks. We observed substantial variation across representations in the ability of LLMs to interpret and generate molecular graphs, with MolJSON consistently outperforming existing formats. On translation tasks, GPT-5 achieved 71.0% accuracy when converting IUPAC names to MolJSON, compared with 43.7% when converting the same inputs to SMILES. For constrained generation, GPT-5 reached 95.3% accuracy generating MolJSON, compared with 76.3% for IUPAC and 64.0% for SMILES. As an input format for shortest-path reasoning, GPT-5 successfully answered 98.5% of questions with MolJSON, compared with 92.2% for SMILES and 82.7% for IUPAC, whilst also using fewer reasoning tokens. We observed systematic errors associated with atom count and ring complexity for SMILES strings and IUPAC names, whereas MolJSON was more robust to these failure modes. Our results show that the choice of molecular representation has a material impact on LLM performance, and that explicit molecular graph schemas, such as MolJSON, are a promising direction for LLM-based systems in chemistry.

📄 PDF Abstract BibTeX arXiv:2605.01822

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

LLaMo: Large Language Model-based Molecular Graph Assistant

2024-10-31 · Jinyoung Park, Minseong Bae, Dohwan Ko, Hyunwoo J. Kim

Large Language Models (LLMs) have demonstrated remarkable generalization and instruction-following capabilities with instruction tuning. The advancements in LLMs and instruction tuning have led to the development of Larg…

Instruction FollowingIUPAC Name PredictionLanguage ModelingLanguage Modelling+4

Two for the Price of One: Integrating Large Language Models to Learn Biophysical Interactions

2025-03-26 · Joseph D. Clark, Tanner J. Dean, Diwakar Shukla

Deep learning models have become fundamental tools in drug design. In particular, large language models trained on biochemical sequences learn feature vectors that guide drug discovery through virtual screening. However,…

Computational EfficiencyDrug DesignDrug DiscoverySpecificity

GIT-Mol: A Multi-modal Large Language Model for Molecular Science with Graph, Image, and Text

2023-08-14 · PengFei Liu, Yiming Ren, Jun Tao, Zhixiang Ren

Large language models have made significant strides in natural language processing, enabling innovative applications in molecular science by processing textual representations of molecules. However, most existing languag…

Drug DiscoveryImage CaptioningLanguage ModelingLanguage Modelling+5

Mol-LLaMA: Towards General Understanding of Molecules in Large Molecular Language Model

2025-02-19 · DongKi Kim, Wonbin Lee, Sung Ju Hwang

Understanding molecules is key to understanding organisms and driving advances in drug discovery, requiring interdisciplinary knowledge across chemistry and biology. Although large molecular language models have achieved…

Drug DiscoveryGeneral KnowledgeLanguage ModelingLanguage Modelling

Implicit Neural Representations of Molecular Vector-Valued Functions

2025-02-15 · Jirka Lhotka, Daniel Probst

Molecules have various computational representations, including numerical descriptors, strings, graphs, point clouds, and surfaces. Each representation method enables the application of various machine learning methodolo…

Decoder