Universal and non-universal text statistics: Clustering coefficient for language identification
In this work we analyze statistical properties of 91 relatively small texts in 7 different languages (Spanish, English, French, German, Turkish, Russian, Icelandic) as well as texts with randomly inserted spaces. Despite the size (around 11260 different words), the well known universal statistical laws -- namely Zipf and Herdan-Heap's laws -- are confirmed, and are in close agreement with results obtained elsewhere. We also construct a word co-occurrence network of each text. While the degree distribution is again universal, we note that the distribution of Clustering Coefficients, which depend strongly on the local structure of networks, can be used to differentiate between languages, as well as to distinguish natural languages from random texts.
Code (0)
등록된 구현이 없습니다.
Tasks
ClusteringLanguage IdentificationSimilar Papers 제목 키워드 기반
Universal statistics of incubation periods and other detection times via diffusion models
We suggest an explanation of typical incubation times statistical features based on the universal behavior of exit times for diffusion models. We give a mathematically rigorous proof of the characteristic right skewness …
Persistence diagrams of random matrices via Morse theory: universality and a new spectral diagnostic
We prove that the persistence diagram of the sublevel set filtration of the quadratic form f(x) = x^T M x restricted to the unit sphere S^{n-1} is analytically determined by the eigenvalues of the symmetric matrix M. By …
Power-law statistics and universal scaling in the absence of criticality
Critical states are sometimes identified experimentally through power-law statistics or universal scaling functions. We show here that such features naturally emerge from networks in self-sustained irregular regimes away…
Neural Scaling Universality: If Exponents Are Fixed, Time to Understand Coefficients
Neural scaling laws describe how pre-training loss decays as power laws with training time, model size, and compute. This position paper argues that the exponents of these power laws are fixed by generic mechanisms: a on…
Developing Universal Dependency Treebanks for Magahi and Braj
In this paper, we discuss the development of treebanks for two low-resourced Indian languages - Magahi and Braj based on the Universal Dependencies framework. The Magahi treebank contains 945 sentences and Braj treebank …