Structured Analysis and Comparison of Alphabets in Historical Handwritten Ciphers
Historical ciphered manuscripts are documents that were typically used in sensitive communications within military and diplomatic contexts or among members of secret societies. These secret messages were concealed by inventing a method of writing employing symbols from diverse sources such as digits, alchemy signs and Latin or Greek characters. When studying a new, unseen cipher, the automatic search and grouping of ciphers with a similar alphabet can aid the scholar in its transcription and cryptanalysis because it indicates a probability that the underlying cipher is similar. In this study, we address this need by proposing the CSI metric, a novel way of comparing pairs of ciphered documents. We assess their effectiveness in an unsupervised clustering scenario utilising visual features, including SIFT, pre-trained learnt embeddings, and OCR descriptors.
Code (0)
등록된 구현이 없습니다.
Tasks
CryptanalysisOptical Character Recognition (OCR)Similar Papers 제목 키워드 기반
Scalable handwritten text recognition system for lexicographic sources of under-resourced languages and alphabets
The paper discusses an approach to decipher large collections of handwritten index cards of historical dictionaries. Our study provides a working solution that reads the cards, and links their lemmas to a searchable list…
Handwritten Text RecognitionHTRTransfer LearningA Few-shot Learning Approach for Historical Ciphered Manuscript Recognition
Encoded (or ciphered) manuscripts are a special type of historical documents that contain encrypted text. The automatic recognition of this kind of documents is challenging because: 1) the cipher alphabet changes from on…
Few-Shot LearningFew-Shot Object DetectionHTRobject-detection+1A Dataset of Inertial Measurement Units for Handwritten English Alphabets
This paper presents an end-to-end methodology for collecting datasets to recognize handwritten English alphabets by utilizing Inertial Measurement Units (IMUs) and leveraging the diversity present in the Indian writing s…
DiversityHandwriting RecognitionOne-shot Compositional Data Generation for Low Resource Handwritten Text Recognition
Low resource Handwritten Text Recognition (HTR) is a hard problem due to the scarce annotated data and the very limited linguistic information (dictionaries and language models). For example, in the case of historical ci…
Handwritten Text RecognitionHTRTransfer Learning using CNN for Handwritten Devanagari Character Recognition
This paper presents an analysis of pre-trained models to recognize handwritten Devanagari alphabets using transfer learning for Deep Convolution Neural Network (DCNN). This research implements AlexNet, DenseNet, Vgg, and…
Transfer Learning