Statistical analysis of word flow among five Indo-European languages
A recent increase in data availability has allowed the possibility to perform different statistical linguistic studies. Here we use the Google Books Ngram dataset to analyze word flow among English, French, German, Italian, and Spanish. We study what we define as ``migrant words'', a type of loanwords that do not change their spelling. We quantify migrant words from one language to another for different decades, and notice that most migrant words can be aggregated in semantic fields and associated to historic events. We also study the statistical properties of accumulated migrant words and their rank dynamics. We propose a measure of use of migrant words that could be used as a proxy of cultural influence. Our methodology is not exempt of caveats, but our results are encouraging to promote further studies in this direction.
Code (0)
등록된 구현이 없습니다.
Similar Papers 제목 키워드 기반
An Analysis of the Differences Among Regional Varieties of Chinese in Malay Archipelago
Chinese features prominently in the Chinese communities located in the nations of Malay Archipelago. In these countries, Chinese has undergone the process of adjustment to the local languages and cultures, which leads to…
Information-theoretical analysis of the statistical dependencies among three variables: Applications to written language
We develop the information-theoretical concepts required to study the statistical dependencies among three variables. Some of such dependencies are pure triple interactions, in the sense that they cannot be explained in …
Anemia, weight, and height among children under five in Peru from 2007 to 2022: A Panel Data analysis
Econometrics in general, and Panel Data methods in particular, are becoming crucial in Public Health Economics and Social Policy analysis. In this discussion paper, we employ a helpful approach of Feasible Generalized Le…
EconometricsOrdinal analysis of lexical patterns
Words are fundamental linguistic units that connect thoughts and things through meaning. However, words do not appear independently in a text sequence. The existence of syntactic rules induces correlations among neighbor…
Time SeriesTime Series AnalysisMimicking Human Process: Text Representation via Latent Semantic Clustering for Classification
Considering that words with different characteristic in the text have different importance for classification, grouping them together separately can strengthen the semantic expression of each part. Thus we propose a new …
ClassificationClusteringGeneral Classification