Subdiffusive semantic evolution in Indo-European languages
How do words change their meaning? Although semantic evolution is driven by a variety of distinct factors, including linguistic, societal, and technological ones, we find that there is one law that holds universally across five major Indo-European languages: that semantic evolution is strongly subdiffusive. Using an automated pipeline of diachronic distributional semantic embedding that controls for underlying symmetries, we show that words follow stochastic trajectories in meaning space with an anomalous diffusion exponent $\alpha= 0.45\pm 0.05$ across languages, in contrast with diffusing particles that follow $\alpha=1$. Randomization methods indicate that preserving temporal correlations in semantic change directions is necessary to recover strongly subdiffusive behavior; however, correlations in change sizes play an important role too. We furthermore show that strong subdiffusion is a robust phenomenon under a wide variety of choices in data analysis and interpretation, such as the choice of fitting an ensemble average of displacements or averaging best-fit exponents of individual word trajectories.
Code (1)
Methods 이 논문이 사용한 방법론
Similar Papers 제목 키워드 기반
Treat the Word As a Whole or Look Inside? Subword Embeddings Model Language Change and Typology
We use a variant of word embedding model that incorporates subword information to characterize the degree of compositionality in lexical semantics. Our models reveal some interesting yet contrastive patterns of long-term…
Proto-Indo-European Lexicon: The Generative Etymological Dictionary of Indo-European Languages
Polish Lexicon-Grammar Development Methodology as an Example for Application to other Languages
In the paper we present our methodology with the intention to propose it as a reference for creating lexicon-grammars. We share our long-term experience gained during research projects (past and on-going) concerning the …
Phylogenetics of Indo-European Language families via an Algebro-Geometric Analysis of their Syntactic Structures
Using Phylogenetic Algebraic Geometry, we analyze computationally the phylogenetic tree of subfamilies of the Indo-European language family, using data of syntactic structures. The two main sources of syntactic data are …
Phonological distances for linguistic typology and the origin of Indo-European languages
We show that short-range phoneme dependencies encode large-scale patterns of linguistic relatedness, with direct implications for quantitative typology and evolutionary linguistics. Specifically, using an information-the…