paper-with-me

홈 › Papers

Direct and indirect evidence of compression of word lengths. Zipf's law of abbreviation revisited

2023-03-17 · Sonia Petrini, Antoni Casas-i-Muñoz, Jordi Cluet-i-Martinell, Mengxue Wang, Chris Bentz, Ramon Ferrer-i-Cancho

Zipf's law of abbreviation, the tendency of more frequent words to be shorter, is one of the most solid candidates for a linguistic universal, in the sense that it has the potential for being exceptionless or with a number of exceptions that is vanishingly small compared to the number of languages on Earth. Since Zipf's pioneering research, this law has been viewed as a manifestation of a universal principle of communication, i.e. the minimization of word lengths, to reduce the effort of communication. Here we revisit the concordance of written language with the law of abbreviation. Crucially, we provide wider evidence that the law holds also in speech (when word length is measured in time), in particular in 46 languages from 14 linguistic families. Agreement with the law of abbreviation provides indirect evidence of compression of languages via the theoretical argument that the law of abbreviation is a prediction of optimal coding. Motivated by the need of direct evidence of compression, we derive a simple formula for a random baseline indicating that word lengths are systematically below chance, across linguistic families and writing systems, and independently of the unit of measurement (length in characters or duration in time). Our work paves the way to measure and compare the degree of optimality of word lengths in languages.

📄 PDF Abstract BibTeX arXiv:2303.10128

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Can Language Models Induce Grammatical Knowledge from Indirect Evidence?

2024-10-08 · Miyu Oba, Yohei Oseki, Akiyo Fukatsu, Akari Haga 외

What kinds of and how much data is necessary for language models to induce grammatical knowledge to judge sentence acceptability? Recent language models still have much room for improvement in their data efficiency compa…

Language AcquisitionSentence

Dependency distance minimization predicts compression

2021-09-18 · Quasy (SyntaxFest) 2021 12 · Ramon Ferrer-i-Cancho, Carlos Gómez-Rodríguez

Dependency distance minimization (DDm) is a well-established principle of word order. It has been predicted theoretically that DDm implies compression, namely the minimization of word lengths. This is a second order pred…

Prediction

Training LLMs over Neurally Compressed Text

2024-04-04 · Brian Lester, Jaehoon Lee, Alex Alemi, Jeffrey Pennington 외

In this paper, we explore the idea of training large language models (LLMs) over highly compressed text. While standard subword tokenizers compress text by a small factor, neural text compressors can achieve much higher …

The optimality of word lengths. Theoretical foundations and an empirical study

2022-08-22 · Sonia Petrini, Antoni Casas-i-Muñoz, Jordi Cluet-i-Martinell, Mengxue Wang 외

Zipf's law of abbreviation, namely the tendency of more frequent words to be shorter, has been viewed as a manifestation of compression, i.e. the minimization of the length of forms -- a universal principle of natural co…

Detecting and Mitigating Indirect Stereotypes in Word Embeddings

2023-05-23 · Erin George, Joyce Chew, Deanna Needell

Societal biases in the usage of words, including harmful stereotypes, are frequently learned by common word embedding methods. These biases manifest not only between a word and an explicit marker of its stereotype, but a…

AttributeWord Embeddings