paper-with-me

홈 › Papers

Similarity of symbol frequency distributions with heavy tails

2015-10-01 · Martin Gerlach, Francesc Font-Clos, Eduardo G. Altmann

Quantifying the similarity between symbolic sequences is a traditional problem in Information Theory which requires comparing the frequencies of symbols in different sequences. In numerous modern applications, ranging from DNA over music to texts, the distribution of symbol frequencies is characterized by heavy-tailed distributions (e.g., Zipf's law). The large number of low-frequency symbols in these distributions poses major difficulties to the estimation of the similarity between sequences, e.g., they hinder an accurate finite-size estimation of entropies. Here we show analytically how the systematic (bias) and statistical (fluctuations) errors in these estimations depend on the sample size~$N$ and on the exponent~$\gamma$ of the heavy-tailed distribution. Our results are valid for the Shannon entropy $(\alpha=1)$, its corresponding similarity measures (e.g., the Jensen-Shanon divergence), and also for measures based on the generalized entropy of order $\alpha$. For small $\alpha$'s, including $\alpha=1$, the errors decay slower than the $1/N$-decay observed in short-tailed distributions. For $\alpha$ larger than a critical value $\alpha^* = 1+1/\gamma \leq 2$, the $1/N$-decay is recovered. We show the practical significance of our results by quantifying the evolution of the English language over the last two centuries using a complete $\alpha$-spectrum of measures. We find that frequent words change more slowly than less frequent words and that $\alpha=2$ provides the most robust measure to quantify language change.

📄 PDF Abstract BibTeX arXiv:1510.00277

Code (0)

등록된 구현이 없습니다.

Tasks

valid

Similar Papers 제목 키워드 기반

Power-law cross-correlations estimation under heavy tails

2016-04-20

We examine the performance of six estimators of the power-law cross-correlations -- the detrended cross-correlation analysis, the detrending moving-average cross-correlation analysis, the height cross-correlation analysi…

Handling Long-tailed Feature Distribution in AdderNets

2021-12-01 · NeurIPS 2021 12 · Minjing Dong, Yunhe Wang, Xinghao Chen, Chang Xu

Adder neural networks (ANNs) are designed for low energy cost which replace expensive multiplications in convolutional neural networks (CNNs) with cheaper additions to yield energy-efficient neural networks and hardware …

Knowledge Distillation

Deep Learning for Improving Numerical Weather Prediction of Heavy Rainfall

2022-03-16 · Journal of Advances in Modeling Earth Systems 2022 3 · Philipp Hess, Niklas Boers

The accurate prediction of rainfall, and in particular of the heaviest rainfall events, remains challenging for numerical weather prediction (NWP) models. This may be due to subgrid-scale parameterizations of processes t…

Deep Learning

Flexible Tails for Normalizing Flows

2024-06-22 · Tennessee Hickling, Dennis Prangle

Normalizing flows are a flexible class of probability distributions, expressed as transformations of a simple base distribution. A limitation of standard normalizing flows is representing distributions with heavy tails, …

Density EstimationVariational Inference

Beyond the Bid-Ask: Strategic Insights into Spread Prediction and the Global Mid-Price Phenomenon

2024-04-17 · Yifan He, Abootaleb Shirvani, Barret Shao, Svetlozar Rachev 외

This research extends the conventional concepts of the bid--ask spread (BAS) and mid-price to include the total market order book bid--ask spread (TMOBBAS) and the global mid-price (GMP). Using high-frequency trading dat…

NavigateNovel Concepts