paper-with-me

홈 › Papers

Improving Chemical Understanding of LLMs via SMILES Parsing

2025-05-22 · Yunhui Jang, Jaehyung Kim, Sungsoo Ahn

Large language models (LLMs) are increasingly recognized as powerful tools for scientific discovery, particularly in molecular science. A fundamental requirement for these models is the ability to accurately understand molecular structures, commonly encoded in the SMILES representation. However, current LLMs struggle to interpret SMILES, even failing to carry out basic tasks such as counting molecular rings. To address this limitation, we introduce CLEANMOL, a novel framework that formulates SMILES parsing into a suite of clean and deterministic tasks explicitly designed to promote graph-level molecular comprehension. These tasks span from subgraph matching to global graph matching, providing structured supervision aligned with molecular structural properties. We construct a molecular pretraining dataset with adaptive difficulty scoring and pre-train open-source LLMs on these tasks. Our results show that CLEANMOL not only enhances structural comprehension but also achieves the best or competes with the baseline on the Mol-Instructions benchmark.

📄 PDF Abstract BibTeX arXiv:2505.16340

Code (0)

등록된 구현이 없습니다.

Tasks

Graph Matchingscientific discovery

Similar Papers 제목 키워드 기반

Empirical Evidence for the Fragment level Understanding on Drug Molecular Structure of LLMs

2024-01-15 · Xiuyuan Hu, Guoqing Liu, Yang Zhao, Hao Zhang

AI for drug discovery has been a research hotspot in recent years, and SMILES-based language models has been increasingly applied in drug molecular design. However, no work has explored whether and how language models un…

Drug DesignDrug Discovery

Can Large Language Models Understand Molecules?

2024-01-05 · Shaghayegh Sadeghi, Alan Bui, Ali Forooghi, Jianguo Lu 외

Purpose: Large Language Models (LLMs) like GPT (Generative Pre-trained Transformer) from OpenAI and LLaMA (Large Language Model Meta AI) from Meta AI are increasingly recognized for their potential in the field of chemin…

Drug DiscoveryLanguage ModellingLarge Language ModelMolecular Property Prediction+3

ChemMLLM: Chemical Multimodal Large Language Model

2025-05-22 · Qian Tan, Dongzhan Zhou, Peng Xia, Wanhao Liu 외

Multimodal large language models (MLLMs) have made impressive progress in many applications in recent years. However, chemical MLLMs that can handle cross-modal understanding and generation remain underexplored. To fill …

Language ModelingLanguage ModellingLarge Language Modelmodel+1

Infinity-Parser2 Technical Report

2026-07-08 · Zuming Huang, Jun Huang, Kexuan Ren, Baode Wang 외 arxiv

We present Infinity-Parser2, a large multimodal model that couples a controllable data-synthesis pipeline with multi-task reinforcement learning for end-to-end document parsing, addressing the persistent scarcity of fait…

Reinforcement Learning

SMILES-Prompting: A Novel Approach to LLM Jailbreak Attacks in Chemical Synthesis

2024-10-21 · Aidan Wong, He Cao, Zijing Liu, Yu Li

The increasing integration of large language models (LLMs) across various fields has heightened concerns about their potential to propagate dangerous information. This paper specifically explores the security vulnerabili…

LLM JailbreakRed Teaming