How Long Is a Piece of String? A Brief Empirical Analysis of Tokenizers
Frontier LLMs are increasingly utilised across academia, society and industry. A commonly used unit for comparing models, their inputs and outputs, and estimating inference pricing is the token. In general, tokens are used as a stable currency, assumed to be broadly consistent across tokenizers and contexts, enabling direct comparisons. However, tokenization varies significantly across models and domains of text, making naive interpretation of token counts problematic. We quantify this variation by providing a comprehensive empirical analysis of tokenization, exploring the compression of sequences to tokens across different distributions of textual data. Our analysis challenges commonly held heuristics about token lengths, finding them to be overly simplistic. We hope the insights of our study add clarity and intuition toward tokenization in contemporary LLMs.
Code (0)
등록된 구현이 없습니다.
Similar Papers 제목 키워드 기반
$π$-augmented pregroups and applications to linguistics
We enrich pregroups with a mapping which allows us to locally apply precyclic permutations to designated substrings. We prove a normalisation theorem for such algebraic structures and briefly formalise some known applica…
The Utility of Text: The Case of Amicus Briefs and the Supreme Court
We explore the idea that authoring a piece of text is an act of maximizing one's expected utility. To make this idea concrete, we consider the societally important decisions of the Supreme Court of the United States. Ext…
counterfactualEfficient Global String Kernel with Random Features: Beyond Counting Substructures
Analysis of large-scale sequential data has been one of the most crucial tasks in areas such as bioinformatics, text, and audio mining. Existing string kernels, however, either (i) rely on local features of short substru…
Quantifying the Complexity of Standard Benchmarking Datasets for Long-Term Human Trajectory Prediction
Methods to quantify the complexity of trajectory datasets are still a missing piece in benchmarking human trajectory prediction models. In order to gain a better understanding of the complexity of trajectory prediction t…
BenchmarkingPredictionQuantizationTrajectory PredictionMultiview Identifiers Enhanced Generative Retrieval
Instead of simply matching a query to pre-existing passages, generative retrieval generates identifier strings of passages as the retrieval target. At a cost, the identifier must be distinctive enough to represent a pass…
Retrieval