Modeling Diachronic Change in Scientific Writing with Information Density
Previous linguistic research on scientific writing has shown that language use in the scientific domain varies considerably in register and style over time. In this paper we investigate the introduction of information theory inspired features to study long term diachronic change on three levels: lexis, part-of-speech and syntax. Our approach is based on distinguishing between sentences from 19th and 20th century scientific abstracts using supervised classification models. To the best of our knowledge, the introduction of information theoretic features to this task is novel. We show that these features outperform more traditional features, such as token or character n-grams, while leading to more compact models. We present a detailed analysis of feature informativeness in order to gain a better understanding of diachronic change on different linguistic levels.
Code (0)
등록된 구현이 없습니다.
Tasks
General ClassificationInformativenessSimilar Papers 제목 키워드 기반
The Royal Society Corpus 6.0: Providing 300+ Years of Scientific Writing for Humanistic Study
We present a new, extended version of the Royal Society Corpus (RSC), a diachronic corpus of scientific English now covering 300+ years of scientific writing (1665--1996). The corpus comprises 47 837 texts, primarily sci…
ArticlesGrammar and Meaning: Analysing the Topology of Diachronic Word Embeddings
The paper showcases the application of word embeddings to change in language use in the domain of science, focusing on the Late Modern English period (17-19th century). Historically, this is the period in which many regi…
ClusteringDiachronic Word EmbeddingsWord EmbeddingsWhat Are LLMs Doing to Scientific Communication? Measuring Changes in Writing Practices and Reading Experience
Has the style of scientific communication changed due to the growing use of large language models in the writing process? We address this question in the domain of Natural Language Processing by leveraging two data resou…
Methods, Data, and Conceptual Change: Reflections from Two Quantitative Diachronic Case Studies
This discussion paper reflects on how quantitative approaches to historical linguistics interact with dataset properties. Drawing on two worked examples, we examine English data using quad-based concept modelling of Earl…
AI-Associated Lexical Shifts Across 34 Languages: Cross-Lingual Convergence and Diachronic Uptake in News Writing
AI-associated lexical shifts have been documented mainly in Scientific English. We extend this work to 34 languages in the WMT News Crawl corpus, refining a split-halves continuation diagnostic that compares GPT-4.1 cont…