paper-with-me

Papers

A statistical test for correspondence of texts to the Zipf-Mandelbrot law

2019-12-25 · Anik Chakrabarty, Mikhail Chebunin, Artyom Kovalevskii, Ilya Pupyshev, Natalia Zakrevskaya, Qianqian Zhou

We analyse correspondence of a text to a simple probabilistic model. The model assumes that the words are selected independently from an infinite dictionary. The probability distribution correspond to the Zipf---Mandelbrot law. We count sequentially the numbers of different words in the text and get the process of the numbers of different words. Then we estimate Zipf---Mandelbrot law parameters using the same sequence and construct an estimate of the expectation of the number of different words in the text. Then we subtract the corresponding values of the estimate from the sequence and normalize along the coordinate axes, obtaining a random process on a segment from 0 to 1. We prove that this process (the empirical text bridge) converges weakly in the uniform metric on $C (0,1)$ to a centered Gaussian process with continuous a.s. paths. We develop and implement an algorithm for approximate calculation of eigenvalues of the covariance function of the limit Gaussian process, and then an algorithm for calculating the probability distribution of the integral of the square of this process. We use the algorithm to analyze uniformity of texts in English, French, Russian and Chinese.

📄 PDF Abstract BibTeX arXiv:1912.11600

Code (0)

등록된 구현이 없습니다.

Methods 이 논문이 사용한 방법론

Gaussian Process Gaussian Processes are non-parametric models for approximating functions. They rely upon a measure of similarity between points (the kernel function) to predict the value for…

Similar Papers 제목 키워드 기반

Is space a word, too?

2017-10-20 · Jake Ryland Williams, Giovanni C. Santia

For words, rank-frequency distributions have long been heralded for adherence to a potentially-universal phenomenon known as Zipf's law. The hypothetical form of this empirical phenomenon was refined by Ben\^{i}ot Mandel…

Compression and the origins of Zipf's law for word frequencies

2016-05-04 · Ramon Ferrer-i-Cancho

Here we sketch a new derivation of Zipf's law for word frequencies based on optimal coding. The structure of the derivation is reminiscent of Mandelbrot's random typing model but it has multiple advantages over random ty…

In narrative texts punctuation marks obey the same statistics as words

2016-04-04 · Andrzej Kulig, Jaroslaw Kwapien, Tomasz Stanisz, Stanislaw Drozdz

From a grammar point of view, the role of punctuation marks in a sentence is formally defined and well understood. In semantic analysis punctuation plays also a crucial role as a method of avoiding ambiguity of the meani…

ArticlesSentence

Words ranking and Hirsch index for identifying the core of the hapaxes in political texts

2020-06-13 · Valerio Ficcadenti, Roy Cerqueti, Marcel Ausloos, Gurjeet Dhesi

This paper deals with a quantitative analysis of the content of official political speeches. We study a set of about one thousand talks pronounced by the US Presidents, ranging from Washington to Trump. In particular, we…

The Surprising Universality of LLM Outputs: A Real-Time Verification Primitive

2026-04-28 · Alex Bogdan, Adrian de Valois-Franklin arxiv

We report a striking statistical regularity in frontier LLM outputs that enables a CPU-only scoring primitive running at 2.6 microseconds per token, with estimated latency up to 100,000$\times$ (five orders of magnitude)…