paper-with-me

홈 › Papers

Structural Pattern Mining in Inka Khipus: Unsupervised Clustering, Provenance Classification, and a Computational Validation of the Santa Valley Match

2026-06-30 · Maria Contreras arxiv

Khipus -- knotted cord devices -- were the primary recording medium of the Inka Empire (c. 1400-1532 CE), yet their system remains undeciphered. We present a reproducible machine-learning pipeline applied to the Open Khipu Repository (OKR), a public database of 619 khipus comprising 54,403 cords and 110,677 knots. We engineer 27 structural features per khipu and apply (i) unsupervised clustering via UMAP and HDBSCAN, recovering three structurally distinct groups (silhouette = 0.769); (ii) supervised provenance classification via gradient boosting, reaching F1 = 0.86 for the Inka Late Horizon imperial style; and (iii) SHAP-based interpretability, which identifies cord twist direction as the dominant structural discriminator of imperial khipus. We further report two findings of methodological interest. First, one cluster is dominated not by a geographic region but by nineteenth-century European museum collections, indicating that colonial acquisition and recording practices are structurally encoded in the corpus. Second, we provide an independent computational verification of the recto/verso (moiety) structure of the six Santa Valley khipus reported by Medrano and Urton (2018), reproducing both the aggregate attachment ratio and the identification of the single mixed specimen--using only the public OKR database, without physical access to the objects. We additionally report a negative result: knot-type sequence order, encoded as n-grams, adds no provenance signal beyond aggregate features. All code and data are openly available.

📄 PDF Abstract BibTeX arXiv:2607.00185

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Exploring Urban Land Use Patterns by Pattern Mining and Unsupervised Learning

2026-03-17 · Zdena Dobesova, Tai Dinh, Pavel Novak arxiv

Urban areas are intricate systems shaped by socioeconomic, environmental, and infrastructural factors, with land use patterns serving as aspects of urban morphology. This paper proposes a novel methodology leveraging fre…

Unsupervised Trajectory Clustering via Adaptive Multi-Kernel-Based Shrinkage

2015-12-01 · ICCV 2015 12 · Hongteng Xu, Yang Zhou, Weiyao Lin, Hongyuan Zha

This paper proposes a shrinkage-based framework for unsupervised trajectory clustering. Facing to the challenges of trajectory clustering, e.g., large variations within a cluster and ambiguities across clusters, we first…

Anomaly DetectionClusteringTrajectory Clustering

Contrastive Entity Linkage: Mining Variational Attributes from Large Catalogs for Entity Linkage

2020-02-14 · AKBC 2020 6 · Varun Embar, Bunyamin Sisman, Hao Wei, Xin Luna Dong 외

Presence of near identical, but distinct, entities called entity variations makes the task of data integration challenging. For example, in the domain of grocery products, variations share the same value for attributes s…

Data Integration

Cross-lingual Joint Entity and Word Embedding to Improve Entity Linking and Parallel Sentence Mining

2019-11-01 · WS 2019 11 · Xiaoman Pan, Thamme Gowda, Heng Ji, Jonathan May 외

Entities, which refer to distinct objects in the real world, can be viewed as language universals and used as effective signals to generate less ambiguous semantic representations and align multiple languages. We propose…

Cross-Lingual Entity LinkingEntity LinkingSentence

Knowledge Discovery using Unsupervised Cognition

2024-09-30 · Alfredo Ibias, Hector Antona, Guillem Ramirez-Miranda, Enric Guinovart

Knowledge discovery is key to understand and interpret a dataset, as well as to find the underlying relationships between its components. Unsupervised Cognition is a novel unsupervised learning algorithm that focus on mo…

Dimensionality Reductionfeature selection