The Royal Society Corpus: From Uncharted Data to Corpus
We present the Royal Society Corpus (RSC) built from the Philosophical Transactions and Proceedings of the Royal Society of London. At present, the corpus contains articles from the first two centuries of the journal (1665―1869) and amounts to around 35 million tokens. The motivation for building the RSC is to investigate the diachronic linguistic development of scientific English. Specifically, we assume that due to specialization, linguistic encodings become more compact over time (Halliday, 1988; Halliday and Martin, 1993), thus creating a specific discourse type characterized by high information density that is functional for expert communication. When building corpora from uncharted material, typically not all relevant meta-data (e.g. author, time, genre) or linguistic data (e.g. sentence/word boundaries, words, parts of speech) is readily available. We present an approach to obtain good quality meta-data and base text data adopting the concept of Agile Software Development.
Code (0)
등록된 구현이 없습니다.
Tasks
ArticlesSentenceSimilar Papers 제목 키워드 기반
The Royal Society Corpus 6.0: Providing 300+ Years of Scientific Writing for Humanistic Study
We present a new, extended version of the Royal Society Corpus (RSC), a diachronic corpus of scientific English now covering 300+ years of scientific writing (1665--1996). The corpus comprises 47 837 texts, primarily sci…
ArticlesThe Making of the Royal Society Corpus
Methods, Data, and Conceptual Change: Reflections from Two Quantitative Diachronic Case Studies
This discussion paper reflects on how quantitative approaches to historical linguistics interact with dataset properties. Drawing on two worked examples, we examine English data using quad-based concept modelling of Earl…
Exploring diachronic syntactic shifts with dependency length: the case of scientific English
We report on an application of universal dependencies for the study of diachronic shifts in syntactic usage patterns. Our focus is on the evolution of Scientific English in the Late Modern English period (ca. 1700-1900).…
Modeling Changing Scientific Concepts with Complex Networks: A Case Study on the Chemical Revolution
While context embeddings produced by LLMs can be used to estimate conceptual change, these representations are often not interpretable nor time-aware. Moreover, bias augmentation in historical data poses a non-trivial ri…