Improving Chemical Understanding of LLMs via SMILES Parsing
Large language models (LLMs) are increasingly recognized as powerful tools for scientific discovery, particularly in molecular science. A fundamental requirement for these models is the ability to accurately understand molecular structures, commonly encoded in the SMILES representation. However, current LLMs struggle to interpret SMILES, even failing to carry out basic tasks such as counting molecular rings. To address this limitation, we introduce CLEANMOL, a novel framework that formulates SMILES parsing into a suite of clean and deterministic tasks explicitly designed to promote graph-level molecular comprehension. These tasks span from subgraph matching to global graph matching, providing structured supervision aligned with molecular structural properties. We construct a molecular pretraining dataset with adaptive difficulty scoring and pre-train open-source LLMs on these tasks. Our results show that CLEANMOL not only enhances structural comprehension but also achieves the best or competes with the baseline on the Mol-Instructions benchmark.
Code (0)
등록된 구현이 없습니다.
Tasks
Graph Matchingscientific discoverySimilar Papers 제목 키워드 기반
Empirical Evidence for the Fragment level Understanding on Drug Molecular Structure of LLMs
AI for drug discovery has been a research hotspot in recent years, and SMILES-based language models has been increasingly applied in drug molecular design. However, no work has explored whether and how language models un…
Drug DesignDrug DiscoveryCan Large Language Models Understand Molecules?
Purpose: Large Language Models (LLMs) like GPT (Generative Pre-trained Transformer) from OpenAI and LLaMA (Large Language Model Meta AI) from Meta AI are increasingly recognized for their potential in the field of chemin…
Drug DiscoveryLanguage ModellingLarge Language ModelMolecular Property Prediction+3ChemMLLM: Chemical Multimodal Large Language Model
Multimodal large language models (MLLMs) have made impressive progress in many applications in recent years. However, chemical MLLMs that can handle cross-modal understanding and generation remain underexplored. To fill …
Language ModelingLanguage ModellingLarge Language Modelmodel+1Infinity-Parser2 Technical Report
We present Infinity-Parser2, a large multimodal model that couples a controllable data-synthesis pipeline with multi-task reinforcement learning for end-to-end document parsing, addressing the persistent scarcity of fait…
Reinforcement LearningSMILES-Prompting: A Novel Approach to LLM Jailbreak Attacks in Chemical Synthesis
The increasing integration of large language models (LLMs) across various fields has heightened concerns about their potential to propagate dangerous information. This paper specifically explores the security vulnerabili…
LLM JailbreakRed Teaming