paper-with-me

홈 › Papers

What Does a Chemical Language Model Know About Molecules?

2026-06-22 · Christian Kenneth, Etowah Adams, Liam Bai, Gerard JP van Westen arxiv

Chemical language models (cLMs) are widely assumed to learn surface-level syntactic patterns rather than learning meaningful molecular semantics. Here, we apply sparse autoencoders (SAEs) to MolFormer, an encoder-only cLM, to mechanistically examine how molecular representations are built across layers. We discover that early layers rely on position-tracking latents to parse molecular grammar, while later layers encode atom-in-substructure and pharmacologically relevant features. Additionally, we show that non-canonical SMILES produce more disruptive representation shifts than invalid SMILES, driven by position-latent disruption propagating across layers. To support further exploration, we develop InterMol, an interactive visualizer for SAE activations on molecular strings and structures.

📄 PDF Abstract BibTeX arXiv:2606.23443

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

What does photon energy tell us about cellphone safety?

2017-04-25

It has been argued that cellphones are safe because a single microwave photon does not have enough energy to break a chemical bond. We show that cellphone technology operates in the classical wave limit, not the single p…

What does it take to bake a cake? The RecipeRef corpus and anaphora resolution in procedural text

2022-05-01 · Findings (ACL) 2022 5 · Biaoyan Fang, Timothy Baldwin, Karin Verspoor

Procedural text contains rich anaphoric phenomena, yet has not received much attention in NLP. To fill this gap, we investigate the textual properties of two types of procedural text, recipes and chemical patents, and ge…

Transfer Learning

Explaining Chemical Toxicity using Missing Features

2020-09-23 · Kar Wai Lim, Bhanushee Sharma, Payel Das, Vijil Chenthamarakshan 외

Chemical toxicity prediction using machine learning is important in drug development to reduce repeated animal and human testing, thus saving cost and time. It is highly recommended that the predictions of computational …

BIG-bench Machine Learning

What does it take to bake a cake? The RecipeRef corpus and anaphora resolution in procedural text

2021-11-16 · ACL ARR November 2021 11 · Anonymous

Procedural text contains rich anaphoric phenomena yet has not received much attention in NLP. To fill this gap, we investigate the textual properties of two types of procedural text, recipes and chemical patents, and gen…

Transfer Learning

Strangers to Themselves: What Language Models Say About Themselves Is Generic

2026-09-09 · Phil Blandfort, Urja Pawar arxiv

Language models can fluently describe how they would behave: whether they would cave to pushback, misuse a tool, or lie under pressure. Is that description actually about the model speaking? We turn self-knowledge into a…