paper-with-me

홈 › Papers

The Functionalizer: Lossless Functional Decomposition for Subword Tokenization

2026-09-18 · Connor Makowski, Willem Guter hf

Standard subword tokenizers either treat every orthographic variation of a word (such as hello, Hello, HELLO, and Héllo) as unrelated vocabulary entries, which fragments the embedding space, or discard this variation through lossy normalization. We present the Functionalizer, a lossless pre-tokenizer framework that factors orthographic and structural variations into a compositional opcode/operand prefix stream before tokenization: a canonical base token (operand) prefixed by parametric transformation operators (opcodes) encoded in the Unicode Private Use Area. We introduce operators covering casing (CAPITALIZE), diacritics (13 dedicated opcodes), and character repetition (REPEAT, MULTIREPEAT), which are fully reversible. Across natural language and code corpora, the Functionalizer enables complete corpus coverage with significantly smaller vocabularies under unconstrained exhaustion conditions, reducing actual vocabulary slot requirements by up to 19.7%. Downstream evaluations on 98M-parameter GPT-2 models show that the Functionalizer improves Python code syntax validity (9.12% vs. 7.70%) while reducing duplicate n-gram repetition in natural language prose. These findings demonstrate that functional decomposition can be an effective mechanism for vocabulary-efficient, structurally aware language modeling, and motivate further validation at production scale.

📄 PDF Abstract BibTeX arXiv:2609.15991

Code (5)

🤗 mrkwanzaa/functionalizer-100M-fineweb-edu-seed1
🤗 mrkwanzaa/functionalizer-100M-fineweb-edu-seed2
🤗 mrkwanzaa/functionalizer-100M-fineweb-edu-seed3
🤗 mrkwanzaa/functionalizer-100M-fineweb-edu-seed4
🤗 mrkwanzaa/functionalizer-100M-fineweb-edu-seed5

Similar Papers 제목 키워드 기반

Lossless Vocabulary Reduction for Auto-Regressive Language Models

2025-10-09 · Daiki Chijiwa, Taku Hasegawa, Kyosuke Nishida, Shin'ya Yamaguchi 외 arxiv

Tokenization -- the process of decomposing a given text into a sequence of subwords called tokens -- is one of the key components in the development of language models. Particularly, auto-regressive language models gener…

Text Generation

Improving Korean NLP Tasks with Linguistically Informed Subword Tokenization and Sub-character Decomposition

2023-11-07 · Taehee Jeon, BongSeok Yang, ChangHwan Kim, Yoonseob Lim

We introduce a morpheme-aware subword tokenization method that utilizes sub-character decomposition to address the challenges of applying Byte Pair Encoding (BPE) to Korean, a language characterized by its rich morpholog…

CoLAComputational EfficiencyMorphological Analysis

Words, Subwords, and Morphemes: What Really Matters in the Surprisal-Reading Time Relationship?

2023-10-26 · Sathvik Nair, Philip Resnik

An important assumption that comes with using LLMs on psycholinguistic data has gone unverified. LLM-based predictions are based on subword tokenization, not decomposition of words into morphemes. Does that matter? We ca…

Effective Context in Transformers: An Analysis of Fragmentation and Tokenization

2026-05-13 · Amirmehdi Jafari Fesharaki, Mohammadamin Rami, Aslan Tchamkerten arxiv

Transformers predict over a representation of a sequence. The same data can be written as bytes, characters, or subword tokens, and these representations may be lossless. Yet, under a fixed context window, they need not …

From Words to Music: A Study of Subword Tokenization Techniques in Symbolic Music Generation

2023-04-18 · Adarsh Kumar, Pedro Sarmento

Subword tokenization has been widely successful in text-based natural language processing (NLP) tasks with Transformer-based models. As Transformer models become increasingly popular in symbolic music-related studies, it…

Music Generation