paper-with-me

홈 › Papers

Where to cut, how deep: BPE and Unigram-LM on chemistry SMILES

2026-07-06 · Hunter Heidenreich hf

Every chemical language model reading SMILES begins with a tokenizer, yet the field has inherited byte-pair encoding (BPE) from natural language with little scrutiny. In natural language, BPE's principal alternative, Unigram-LM, is known to build structurally different vocabularies. Whether that contrast survives in chemistry was open. We report a controlled comparison of BPE and Unigram-LM over a fixed 165-token chemistry base, at the small vocabulary sizes where token embeddings are learnable, across three corpus typologies (diverse, drug-like, natural-products) and both pre-tokenization boundary policies. The two do not converge. In all 22 matched conditions they build near-disjoint subword vocabularies: cross-algorithm Jaccard overlap on the learned pieces never exceeds 0.161, and at most 0.05 once weighted toward the high-frequency pieces a model updates most. Unigram-LM also segments held-out molecules into 29-41% more tokens; the arms largely agree on where to cut but not how deeply, so BPE's segmentation is a strict coarsening of Unigram-LM's on 80-99% of molecules. The separation holds across corpus, boundary, and vocabulary size, persisting even at eight times that scale. The subword algorithm is therefore a modeling decision, not a free default. The study trains no language models.

📄 PDF Abstract BibTeX arXiv:2607.05691

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Beyond Chemical 1D knowledge using Transformers

2020-10-02 · Ruud Van Deursen, Igor V. Tetko, Guillaume Godin

In the present paper we evaluated efficiency of the recent Transformer-CNN models to predict target properties based on the augmented stereochemical SMILES. We selected a well-known Cliff activity dataset as well as a Di…

Molecular Representations for Large Language Models

2026-05-03 · Nicholas T. Runcie, Fergus Imrie, Charlotte M. Deane arxiv

Large Language Models (LLMs) are increasingly being used to support scientific discovery. In chemistry, tasks such as reaction prediction and structure elucidation require reasoning about the structures of molecules. As …

MinerU.Chem: A High-Precision System for Optical Chemical Structure and Reaction Recognition

2026-08-04 · Haote Yang, Jiang Wu, Jingchao Wang, Xingjian Wei 외 arxiv

In organic chemistry papers and patents, molecular structures, reaction schemes, and experimental conditions are often presented as molecular structure depictions, reaction diagrams, and complex tables or figures. Such i…

Molecular Property Prediction

Rethinking Molecular Text Representations for LLMs: An Empirical Study

2026-06-02 · Arun Raja, Garrett M. Morris, Kian Ming A. Chai arxiv

Large language models (LLMs) are increasingly used for molecular tasks, but it remains unclear which molecular representation to use. We present a systematic benchmark evaluating LLM molecular competence across nine repr…

Bi-semantic Chemical Embedder for Joint Representation Learning of SMILES and Natural Language

2026-08-04 · David Ming Segura, Jeremy Goumaz, Joshua W. Sin, Bojana Ranković 외 arxiv

Transformer models have revolutionized natural language processing (NLP), and text-based molecular representations like SMILES have successfully extended these architectures to chemistry. However, domain-adaptive pre-tra…

Molecular Property PredictionRepresentation LearningContrastive Learning