paper-with-me

홈 › Papers

Universal and non-universal text statistics: Clustering coefficient for language identification

2019-11-18 · Diego Espitia, Hernán Larralde

In this work we analyze statistical properties of 91 relatively small texts in 7 different languages (Spanish, English, French, German, Turkish, Russian, Icelandic) as well as texts with randomly inserted spaces. Despite the size (around 11260 different words), the well known universal statistical laws -- namely Zipf and Herdan-Heap's laws -- are confirmed, and are in close agreement with results obtained elsewhere. We also construct a word co-occurrence network of each text. While the degree distribution is again universal, we note that the distribution of Clustering Coefficients, which depend strongly on the local structure of networks, can be used to differentiate between languages, as well as to distinguish natural languages from random texts.

📄 PDF Abstract BibTeX arXiv:1911.08915

Code (0)

등록된 구현이 없습니다.

Tasks

ClusteringLanguage Identification

Similar Papers 제목 키워드 기반

Universal statistics of incubation periods and other detection times via diffusion models

2018-04-16

We suggest an explanation of typical incubation times statistical features based on the universal behavior of exit times for diffusion models. We give a mathematically rigorous proof of the characteristic right skewness …

Persistence diagrams of random matrices via Morse theory: universality and a new spectral diagnostic

2026-03-29 · Matthew Loftus arxiv

We prove that the persistence diagram of the sublevel set filtration of the quadratic form f(x) = x^T M x restricted to the unit sphere S^{n-1} is analytically determined by the eigenvalues of the symmetric matrix M. By …

Power-law statistics and universal scaling in the absence of criticality

2015-03-27 · Jonathan Touboul, Alain Destexhe

Critical states are sometimes identified experimentally through power-law statistics or universal scaling functions. We show here that such features naturally emerge from networks in self-sustained irregular regimes away…

Neural Scaling Universality: If Exponents Are Fixed, Time to Understand Coefficients

2026-06-23 · Yizhou Liu, Jeff Gore arxiv

Neural scaling laws describe how pre-training loss decays as power laws with training time, model size, and compute. This position paper argues that the exponents of these power laws are fixed by generic mechanisms: a on…

Developing Universal Dependency Treebanks for Magahi and Braj

2022-04-26 · Mohit Raj, Shyam Ratan, Deepak Alok, Ritesh Kumar 외

In this paper, we discuss the development of treebanks for two low-resourced Indian languages - Magahi and Braj based on the Universal Dependencies framework. The Magahi treebank contains 945 sentences and Braj treebank …