paper-with-me

홈 › Papers

mCLM: A Function-Infused and Synthesis-Friendly Modular Chemical Language Model

2025-05-18 · Carl Edwards, Chi Han, Gawon Lee, Thao Nguyen, Bowen Jin, Chetan Kumar Prasad, Sara Szymkuć, Bartosz A. Grzybowski, Ying Diao, Jiawei Han, Ge Liu, Hao Peng, Martin D. Burke, Heng Ji

Despite their ability to understand chemical knowledge and accurately generate sequential representations, large language models (LLMs) remain limited in their capacity to propose novel molecules with drug-like properties. In addition, the molecules that LLMs propose can often be challenging to make in the lab. To more effectively enable the discovery of functional small molecules, LLMs need to learn a molecular language. However, LLMs are currently limited by encoding molecules from atoms. In this paper, we argue that just like tokenizing texts into (sub-)word tokens instead of characters, molecules should be decomposed and reassembled at the level of functional building blocks, i.e., parts of molecules that bring unique functions and serve as effective building blocks for real-world automated laboratory synthesis. This motivates us to propose mCLM, a modular Chemical-Language Model tokenizing molecules into building blocks and learning a bilingual language model of both natural language descriptions of functions and molecule building blocks. By reasoning on such functional building blocks, mCLM guarantees to generate efficiently synthesizable molecules thanks to recent progress in block-based chemistry, while also improving the functions of molecules in a principled manner. In experiments on 430 FDA-approved drugs, we find mCLM capable of significantly improving 5 out of 6 chemical functions critical to determining drug potentials. More importantly, mCLM can reason on multiple functions and improve the FDA-rejected drugs (``fallen angels'') over multiple iterations to greatly improve their shortcomings.

📄 PDF Abstract BibTeX arXiv:2505.12565

Code (0)

등록된 구현이 없습니다.

Tasks

Language ModelingLanguage Modelling

Similar Papers 제목 키워드 기반

MCLMR: A Model-Agnostic Causal Learning Framework for Multi-Behavior Recommendation

2026-03-26 · Ranxu Zhang, Junjie Meng, Ying Sun, Ziqi Xu 외 arxiv

Multi-Behavior Recommendation (MBR) leverages multiple user interaction types (e.g., views, clicks, purchases) to enrich preference modeling and alleviate data sparsity issues in traditional single-behavior approaches. H…

Contrastive Learning

SMCLM: Semantically Meaningful Causal Language Modeling for Autoregressive Paraphrase Generation

2025-07-04 · Michał Perełkiewicz, Sławomir Dadas, Rafał Poświata arxiv

This article introduces semantically meaningful causal language modeling (SMCLM), a selfsupervised method of training autoregressive models to generate semantically equivalent text. Our approach involves using semantical…

Paraphrase Generation

Verilog-Evolve: Feedback-Driven and Skill-Evolving Verilog Generation

2026-05-26 · Zehua Pei, Hui-Ling Zhen, Yu Zhang, Sinno Jialin Pan 외 arxiv

Large language models (LLMs) have improved Verilog generation from natural-language specifications, but most pipelines still treat generation as isolated sampling followed by functional checking. This is insufficient for…

Linguistic Generalizability of Test-Time Scaling in Mathematical Reasoning

2025-02-24 · Guijin Son, Jiwoo Hong, Hyunwoo Ko, James Thorne

Scaling pre-training compute has proven effective for achieving mulitlinguality, but does the same hold for test-time scaling? In this work, we introduce MCLM, a multilingual math benchmark featuring competition-level pr…

MathMathematical Reasoning

Fluctuation without dissipation: Microcanonical Langevin Monte Carlo

2023-03-31 · Jakob Robnik, Uroš Seljak

Stochastic sampling algorithms such as Langevin Monte Carlo are inspired by physical systems in a heat bath. Their equilibrium distribution is the canonical ensemble given by a prescribed target distribution, so they mus…