Using Comparable Collections of Historical Texts for Building a Diachronic Dictionary for Spelling Normalization
Code (0)
등록된 구현이 없습니다.
Similar Papers 제목 키워드 기반
Topic modelling discourse dynamics in historical newspapers
This paper addresses methodological issues in diachronic data analysis for historical research. We apply two families of topic models (LDA and DTM) on a relatively large set of historical newspapers, with the aim of capt…
Topic ModelsOne-to-X analogical reasoning on word embeddings: a case for diachronic armed conflict prediction from news texts
We extend the well-known word analogy task to a one-to-X formulation, including one-to-none cases, when no correct answer exists. The task is cast as a relation discovery problem and applied to historical armed conflicts…
Word EmbeddingsSiDiaC: Sinhala Diachronic Corpus
SiDiaC, the first comprehensive Sinhala Diachronic Corpus, covers a historical span from the 5th to the 20th century CE. SiDiaC comprises 58k words across 46 literary works, annotated carefully based on the written date,…
Document AIA Neural Model for Part-of-Speech Tagging in Historical Texts
Historical texts are challenging for natural language processing because they differ linguistically from modern texts and because of their lack of orthographical and grammatical standardisation. We use a character-level …
Part-Of-Speech TaggingPOSPOS TaggingNederlab: Towards a Single Portal and Research Environment for Diachronic Dutch Text Corpora
The Nederlab project aims to bring together all digitized texts relevant to the Dutch national heritage, the history of the Dutch language and culture (circa 800 {--} present) in one user friendly and tool enriched open …
Cultural Vocal Bursts Intensity Prediction