paper-with-me

홈 › Papers

Random Text, Zipf's Law, Critical Length,and Implications for Large Language Models

2025-11-14 · Vladimir Berman arxiv

We study a deliberately simple, fully non-linguistic model of text: a sequence of independent draws from a finite alphabet of letters plus a single space symbol. A word is defined as a maximal block of non-space symbols. Within this symbol-level framework, which assumes no morphology, syntax, or semantics, we derive several structural results. First, word lengths follow a geometric distribution governed solely by the probability of the space symbol. Second, the expected number of words of a given length, and the expected number of distinct words of that length, admit closed-form expressions based on a coupon-collector argument. This yields a critical word length k* at which word types transition from appearing many times on average to appearing at most once. Third, combining the exponential growth of the number of possible strings of length k with the exponential decay of the probability of each string, we obtain a Zipf-type rank-frequency law p(r) proportional to r^{-alpha}, with an exponent determined explicitly by the alphabet size and the space probability. Our contribution is twofold. Mathematically, we give a unified derivation linking word lengths, vocabulary growth, critical length, and rank-frequency structure in a single explicit model. Conceptually, we argue that this provides a structurally grounded null model for both natural-language word statistics and token statistics in large language models. The results show that Zipf-like patterns can arise purely from combinatorics and segmentation, without optimization principles or linguistic organization, and help clarify which phenomena require deeper explanation beyond random-text structure.

📄 PDF Abstract BibTeX arXiv:2511.17575

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Optimal coding and the origins of Zipfian laws

2019-06-04 · Ramon Ferrer-i-Cancho, Christian Bentz, Caio Seguin

The problem of compression in standard information theory consists of assigning codes as short as possible to numbers. Here we consider the problem of optimal coding -- under an arbitrary coding scheme -- and show that i…

Zipf's law holds for phrases, not words

2014-06-19 · Jake Ryland Williams, Paul R. Lessard, Suma Desu, Eric Clark 외

With Zipf's law being originally and most famously observed for word frequency, it is surprisingly limited in its applicability to human language, holding over no more than three to four orders of magnitude before hittin…

The Morphemic Origin of Zipf's Law: A Factorized Combinatorial Framework

2025-12-13 · Vladimir Berman arxiv

We present a simple structure based model of how words are formed from morphemes. The model explains two major empirical facts: the typical distribution of word lengths and the appearance of Zipf like rank frequency curv…

Direct and indirect evidence of compression of word lengths. Zipf's law of abbreviation revisited

2023-03-17 · Sonia Petrini, Antoni Casas-i-Muñoz, Jordi Cluet-i-Martinell, Mengxue Wang 외

Zipf's law of abbreviation, the tendency of more frequent words to be shorter, is one of the most solid candidates for a linguistic universal, in the sense that it has the potential for being exceptionless or with a numb…

Revisiting the Optimality of Word Lengths

2023-12-06 · Tiago Pimentel, Clara Meister, Ethan Gotlieb Wilcox, Kyle Mahowald 외

Zipf (1935) posited that wordforms are optimized to minimize utterances' communicative costs. Under the assumption that cost is given by an utterance's length, he supported this claim by showing that words' lengths are i…