paper-with-me

홈 › Papers

Pre-trained Models Perform the Best When Token Distributions Follow Zipf's Law

2025-07-30 · Yanjin He, Qingkai Zeng, Meng Jiang arxiv

Tokenization is a fundamental step in natural language processing (NLP) and other sequence modeling domains, where the choice of vocabulary size significantly impacts model performance. Despite its importance, selecting an optimal vocabulary size remains underexplored, typically relying on heuristics or dataset-specific choices. In this work, we propose a principled method for determining the vocabulary size by analyzing token frequency distributions through Zipf's law. We show that downstream task performance correlates with how closely token distributions follow power-law behavior, and that aligning with Zipfian scaling improves both model efficiency and effectiveness. Extensive experiments across NLP, genomics, and chemistry demonstrate that models consistently achieve peak performance when the token distribution closely adheres to Zipf's law, establishing Zipfian alignment as a robust and generalizable criterion for vocabulary size selection.

📄 PDF Abstract BibTeX arXiv:2507.22543

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Next-Token Prediction and Regret Minimization

2026-03-30 · Mehryar Mohri, Clayton Sanford, Jon Schneider, Kiran Vodrahalli 외 arxiv

We consider the question of how to employ next-token prediction algorithms in adversarial online decision-making environments. Specifically, if we train a next-token prediction model on a distribution $\mathcal{D}$ over …

On Next-Token Prediction in LLMs: How End Goals Determine the Consistency of Decoding Algorithms

2025-05-16 · Jacob Trauger, Ambuj Tewari

Probabilistic next-token prediction trained using cross-entropy loss is the basis of most large language models. Given a sequence of previous values, next-token prediction assigns a probability to each possible next valu…

Information RetrievalPrediction

Provable Generalization of SGD-trained Neural Networks of Any Width in the Presence of Adversarial Label Noise

2021-01-04 · Spencer Frei, Yuan Cao, Quanquan Gu

We consider a one-hidden-layer leaky ReLU network of arbitrary width trained by stochastic gradient descent (SGD) following an arbitrary initialization. We prove that SGD produces neural networks that have classification…

Trained Transformers Learn Linear Models In-Context

2023-06-16 · Ruiqi Zhang, Spencer Frei, Peter L. Bartlett

Attention-based neural networks such as transformers have demonstrated a remarkable ability to exhibit in-context learning (ICL): Given a short prompt sequence of tokens from an unseen task, they can formulate relevant p…

In-Context Learningregression

Can Released LLM Vocabularies Support Token-Level Estimation of Hidden Corpora?

2026-08-11 · Qingjie Zhang, Xingzhang Ren, Zixuan Chen, Jinfeng Li 외 arxiv

Pretraining corpus composition shapes LLM capabilities, but it often remains hidden even when model weights are released. Prior work has inferred corpus mixtures or traced specific token groups from released tokenizer vo…