paper-with-me

홈 › Papers

The Royal Society Corpus: From Uncharted Data to Corpus

2016-05-01 · LREC 2016 5 · Hannah Kermes, Stefania Degaetano-Ortlieb, Ashraf Khamis, J{\"o}rg Knappen, Elke Teich

We present the Royal Society Corpus (RSC) built from the Philosophical Transactions and Proceedings of the Royal Society of London. At present, the corpus contains articles from the first two centuries of the journal (1665―1869) and amounts to around 35 million tokens. The motivation for building the RSC is to investigate the diachronic linguistic development of scientific English. Specifically, we assume that due to specialization, linguistic encodings become more compact over time (Halliday, 1988; Halliday and Martin, 1993), thus creating a specific discourse type characterized by high information density that is functional for expert communication. When building corpora from uncharted material, typically not all relevant meta-data (e.g. author, time, genre) or linguistic data (e.g. sentence/word boundaries, words, parts of speech) is readily available. We present an approach to obtain good quality meta-data and base text data adopting the concept of Agile Software Development.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Tasks

ArticlesSentence

Similar Papers 제목 키워드 기반

The Royal Society Corpus 6.0: Providing 300+ Years of Scientific Writing for Humanistic Study

2020-05-01 · LREC 2020 5 · Stefan Fischer, J{\"o}rg Knappen, Katrin Menzel, Elke Teich

We present a new, extended version of the Royal Society Corpus (RSC), a diachronic corpus of scientific English now covering 300+ years of scientific writing (1665--1996). The corpus comprises 47 837 texts, primarily sci…

Articles

The Making of the Royal Society Corpus

2017-05-01 · WS 2017 5 · J{\"o}rg Knappen, Stefan Fischer, Hannah Kermes, Elke Teich 외
Optical Character Recognition (OCR)Part-Of-Speech TaggingWord Embeddings

Methods, Data, and Conceptual Change: Reflections from Two Quantitative Diachronic Case Studies

2026-05-03 · Catherine Wong, Bach Phan-Tat, Susan Fitzmaurice arxiv

This discussion paper reflects on how quantitative approaches to historical linguistics interact with dataset properties. Drawing on two worked examples, we examine English data using quad-based concept modelling of Earl…

Exploring diachronic syntactic shifts with dependency length: the case of scientific English

2020-12-01 · UDW (COLING) 2020 12 · Tom S Juzek, Marie-Pauline Krielke, Elke Teich

We report on an application of universal dependencies for the study of diachronic shifts in syntactic usage patterns. Our focus is on the evolution of Scientific English in the Late Modern English period (ca. 1700-1900).…

Modeling Changing Scientific Concepts with Complex Networks: A Case Study on the Chemical Revolution

2026-03-18 · Sofía Aguilar-Valdez, Stefania Degaetano-Ortlieb arxiv

While context embeddings produced by LLMs can be used to estimate conceptual change, these representations are often not interpretable nor time-aware. Moreover, bias augmentation in historical data poses a non-trivial ri…