paper-with-me

홈 › Papers

Range-Limited Heaps' Law for Functional DNA Words in the Human Genome

2024-05-22 · Wentian Li, Yannis Almirantis, Astero Provata

Heaps' or Herdan's law is a linguistic law describing the relationship between the vocabulary/dictionary size (type) and word counts (token) to be a power-law function. Its existence in genomes with certain definition of DNA words is unclear partly because the dictionary size in genome could be much smaller than that in a human language. We define a DNA word as a coding region in a genome that codes for a protein domain. Using human chromosomes and chromosome arms as individual samples, we establish the existence of Heaps' law in the human genome within limited range. Our definition of words in a genomic or proteomic context is different from other definitions such as over-represented k-mers which are much shorter in length. Although an approximate power-law distribution of protein domain sizes due to gene duplication and the related Zipf's law is well known, their translation to the Heaps' law in DNA words is not automatic. Several other animal genomes are shown herein also to exhibit range-limited Heaps' law with our definition of DNA words, though with various exponents. When tokens were randomly sampled and sample sizes reach to the maximum level, a deviation from the Heaps' law was observed, but a quadratic regression in log-log type-token plot fits the data perfectly. Investigation of type-token plot and its regression coefficients could provide an alternative narrative of reusage and redundancy of protein domains as well as creation of new protein domains from a linguistic perspective.

📄 PDF Abstract BibTeX arXiv:2405.13825

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Heaps' Law in GPT-Neo Large Language Model Emulated Corpora

2023-11-10 · Uyen Lai, Gurjit S. Randhawa, Paul Sheridan

Heaps' law is an empirical relation in text analysis that predicts vocabulary growth as a function of corpus size. While this law has been validated in diverse human-authored text corpora, its applicability to large lang…

Language ModelingLanguage ModellingLarge Language Model

Heaps' law and Heaps functions in tagged texts: Evidences of their linguistic relevance

2020-01-07 · Andrés Chacoma, Damián H. Zanette

We study the relationship between vocabulary size and text length in a corpus of $75$ literary works in English, authored by six writers, distinguishing between the contributions of three grammatical classes (or ``tags,'…

TAG

Scaling laws in human speech, decreasing emergence of new words and a generalized model

2014-12-16 · Ruokuang Lin, Qianli D. Y. Ma, Chunhua Bian

Human language, as a typical complex system, its organization and evolution is an attractive topic for both physical and cultural researchers. In this paper, we present the first exhaustive analysis of the text organizat…

Corrections of Zipf's and Heaps' Laws Derived from Hapax Rate Models

2023-07-24 · Łukasz Dębowski

The article introduces corrections to Zipf's and Heaps' laws based on systematic models of the proportion of hapaxes, i.e., words that occur once. The derivation rests on two assumptions: The first one is the standard ur…

Statistical laws and linguistics differ in naturalistic video and fictional conversations

2025-12-19 · Ashley M. A. Fehr, Calla G. Beauregard, Julia Witte Zimmerman, Katie Ekström 외 arxiv

Conversation is a cornerstone of social connection and is linked to well-being outcomes. Conversations vary widely in type with some portion generating complex, dynamic stories. One approach to studying how conversations…