paper-with-me

홈 › Papers

Character Entropy in Modern and Historical Texts: Comparison Metrics for an Undeciphered Manuscript

2020-10-28 · Luke Lindemann, Claire Bowern

This paper outlines the creation of three corpora for multilingual comparison and analysis of the Voynich manuscript: a corpus of Voynich texts partitioned by Currier language, scribal hand, and transcription system, a corpus of 294 language samples compiled from Wikipedia, and a corpus of eighteen transcribed historical texts in eight languages. These corpora will be utilized in subsequent work by the Voynich Working Group at Yale University. We demonstrate the utility of these corpora for studying characteristics of the Voynich script and language, with an analysis of conditional character entropy in Voynichese. We discuss the interaction between character entropy and language, script size and type, glyph compositionality, scribal conventions and abbreviations, positional character variants, and bigram frequency. This analysis characterizes the interaction between script compositionality, character size, and predictability. We show that substantial manipulations of glyph composition are not sufficient to align conditional entropy levels with natural languages. The unusually predictable nature of the Voynichese script is not attributable to a particular script or transcription system, underlying language, or substitution cipher. Voynichese is distinct from every comparison text in our corpora because character placement is highly constrained within the word, and this may indicate the loss of phonemic distinctions from the underlying language.

📄 PDF Abstract BibTeX arXiv:2010.14697

Code (1)

chirila/Voynich-public 공식 구현

Similar Papers 제목 키워드 기반

Automatic Orality Identification in Historical Texts

2020-05-01 · LREC 2020 5 · Katrin Ortmann, Stefanie Dipper

Independently of the medial representation (written/spoken), language can exhibit characteristics of conceptual orality or literacy, which mainly manifest themselves on the lexical or syntactic level. In this paper we ai…

Sentence

A Neural Model for Part-of-Speech Tagging in Historical Texts

2016-12-01 · COLING 2016 12 · Christian Hardmeier

Historical texts are challenging for natural language processing because they differ linguistically from modern texts and because of their lack of orthographical and grammatical standardisation. We use a character-level …

Part-Of-Speech TaggingPOSPOS Tagging

Automatic Normalisation of Early Modern French

2022-06-01 · LREC 2022 6 · Rachel Bawden, Jonathan Poinhos, Eleni Kogkitsidou, Philippe Gambette 외

Spelling normalisation is a useful step in the study and analysis of historical language texts, whether it is manual analysis by experts or automatic analysis using downstream natural language processing (NLP) tools. Not…

Deep learning enables urban change profiling through alignment of historical maps

2026-02-02 · Sidi Wu, Yizi Chen, Maurizio Gribaudi, Konrad Schindler 외 arxiv

Prior to modern Earth observation technologies, historical maps provide a unique record of long-term urban transformation and offer a lens on the evolving identity of cities. However, extracting consistent and fine-grain…

Object Detection

InteChar: A Unified Oracle Bone Character List for Ancient Chinese Language Modeling

2025-08-12 · Xiaolei Diao, Zhihan Zhou, Lida Shi, Ting Wang 외 arxiv

Constructing historical language models (LMs) plays a crucial role in aiding archaeological provenance studies and understanding ancient cultures. However, existing resources present major challenges for training effecti…

Data Augmentation