paper-with-me

홈 › Papers

Entropy and type-token ratio in gigaword corpora

2024-11-15 · Pablo Rosillo-Rodes, Maxi San Miguel, David Sanchez

There are different ways of measuring diversity in complex systems. In particular, in language, lexical diversity is characterized in terms of the type-token ratio and the word entropy. We here investigate both diversity metrics in six massive linguistic datasets in English, Spanish, and Turkish, consisting of books, news articles, and tweets. These gigaword corpora correspond to languages with distinct morphological features and differ in registers and genres, thus constituting a varied testbed for a quantitative approach to lexical diversity. We unveil an empirical functional relation between entropy and type-token ratio of texts of a given corpus and language, which is a consequence of the statistical laws observed in natural language. Further, in the limit of large text lengths we find an analytical expression for this relation relying on both Zipf and Heaps laws that agrees with our empirical findings.

📄 PDF Abstract BibTeX arXiv:2411.10227

Code (0)

등록된 구현이 없습니다.

Tasks

ArticlesDiversityRelation

Similar Papers 제목 키워드 기반

Sparse Non-negative Matrix Language Modeling

2016-01-01 · TACL 2016 1 · Joris Pelemans, Noam Shazeer, Ciprian Chelba

We present Sparse Non-negative Matrix (SNM) estimation, a novel probability estimation technique for language modeling that can efficiently incorporate arbitrary features. We evaluate SNM language models on two corpora: …

Automatic Speech Recognition (ASR)Language ModelingLanguage ModellingSentence+1

Corpora Compared: The Case of the Swedish Gigaword & Wikipedia Corpora

2020-11-06 · Tosin P. Adewumi, Foteini Liwicki, Marcus Liwicki

In this work, we show that the difference in performance of embeddings from differently sourced data for a given language can be due to other factors besides data size. Natural language processing (NLP) tasks usually per…

The Danish Gigaword Corpus

2021-05-01 · NoDaLiDa 2021 5 · Leon Strømberg-Derczynski, Manuel Ciosici, Rebekah Baglini, Morten H. Christiansen 외

Danish language technology has been hindered by a lack of broad-coverage corpora at the scale modern NLP prefers. This paper describes the Danish Gigaword Corpus, the result of a focused effort to provide a diverse and f…

The Danish Gigaword Project

2020-05-07 · Leon Strømberg-Derczynski, Manuel R. Ciosici, Rebekah Baglini, Morten H. Christiansen 외

Danish language technology has been hindered by a lack of broad-coverage corpora at the scale modern NLP prefers. This paper describes the Danish Gigaword Corpus, the result of a focused effort to provide a diverse and f…

Geographically-Balanced Gigaword Corpora for 50 Language Varieties

2020-05-01 · LREC 2020 5 · Jonathan Dunn, Ben Adams

While text corpora have been steadily increasing in overall size, even very large corpora are not designed to represent global population demographics. For example, recent work has shown that existing English gigaword co…

Word Embeddings