paper-with-me

홈 › Papers

Tokenizer Choice For LLM Training: Negligible or Crucial?

2023-10-12 · Mehdi Ali, Michael Fromm, Klaudia Thellmann, Richard Rutmann, Max Lübbering, Johannes Leveling, Katrin Klug, Jan Ebert, Niclas Doll, Jasper Schulze Buschhoff, Charvi Jain, Alexander Arno Weber, Lena Jurkschat, Hammam Abdelwahab, Chelsea John, Pedro Ortiz Suarez, Malte Ostendorff, Samuel Weinbach, Rafet Sifa, Stefan Kesselheim, Nicolas Flores-Herr

The recent success of Large Language Models (LLMs) has been predominantly driven by curating the training dataset composition, scaling of model architectures and dataset sizes and advancements in pretraining objectives, leaving tokenizer influence as a blind spot. Shedding light on this underexplored area, we conduct a comprehensive study on the influence of tokenizer choice on LLM downstream performance by training 24 mono- and multilingual LLMs at a 2.6B parameter scale, ablating different tokenizer algorithms and parameterizations. Our studies highlight that the tokenizer choice can significantly impact the model's downstream performance and training costs. In particular, we find that the common tokenizer evaluation metrics fertility and parity are not always predictive of model downstream performance, rendering these metrics a questionable proxy for the model's downstream performance. Furthermore, we show that multilingual tokenizers trained on the five most frequent European languages require vocabulary size increases of factor three in comparison to English. While English-centric tokenizers have been applied to the training of multi-lingual LLMs in the past, we find that this approach results in a severe downstream performance degradation and additional training costs of up to 68%, due to an inefficient tokenization vocabulary.

📄 PDF Abstract BibTeX arXiv:2310.08754

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Beyond Text Compression: Evaluating Tokenizers Across Scales

2025-06-03 · Jonas F. Lotz, António V. Lopes, Stephan Peitz, Hendra Setiawan 외

The choice of tokenizer can profoundly impact language model performance, yet accessible and reliable evaluations of tokenizer quality remain an open challenge. Inspired by scaling consistency, we show that smaller model…

Language ModelingLanguage ModellingText Compression

MUTANT: A Recipe for Multilingual Tokenizer Design

2025-11-05 · Souvik Rana, Arul Menezes, Ashish Kulkarni, Chandra Khatri 외 arxiv

Tokenizers play a crucial role in determining the performance, training efficiency, and the inference cost of Large Language Models (LLMs). Designing effective tokenizers for multilingual LLMs is particularly challenging…

How Good is Your Tokenizer? On the Monolingual Performance of Multilingual Language Models

2020-12-31 · ACL 2021 5 · Phillip Rust, Jonas Pfeiffer, Ivan Vulić, Sebastian Ruder 외

In this work, we provide a systematic and comprehensive empirical comparison of pretrained multilingual language models versus their monolingual counterparts with regard to their monolingual task performance. We study a …

Pretrained Multilingual Language Models

What Makes for Good Tokenizers in Vision Transformer?

2022-12-21 · Shengju Qian, Yi Zhu, Wenbo Li, Mu Li 외

The architecture of transformers, which recently witness booming applications in vision tasks, has pivoted against the widespread convolutional paradigm. Relying on the tokenization process that splits inputs into multip…

TokEval: A Tokenizer Evaluation Suite

2026-08-18 · Clara Meister arxiv

Language model tokenizers are typically selected with minimal evaluation, despite the fact that their design choices directly impact model capabilities. This can be partly attributed to a limited understanding of which t…

Mathematical ReasoningCode Generation