Similarity Dependent Chinese Restaurant Process for Cognate Identification in Multilingual Wordlists
We present and evaluate two similarity dependent Chinese Restaurant Process (sd-CRP) algorithms at the task of automated cognate detection. The sd-CRP clustering algorithms do not require any predefined threshold for detecting cognate sets in a multilingual word list. We evaluate the performance of the algorithms on six language families (more than 750 languages) and find that both the sd-CRP variants performs as well as InfoMap and better than UPGMA at the task of inferring cognate clusters. The algorithms presented in this paper are family agnostic and can be applied to any linguistically under-studied language family.
Code (0)
등록된 구현이 없습니다.
Tasks
ClusteringLanguage IdentificationSimilar Papers 제목 키워드 기반
Chinese Restaurant Process for cognate clustering: A threshold free approach
In this paper, we introduce a threshold free approach, motivated from Chinese Restaurant Process, for the purpose of cognate clustering. We show that our approach yields similar results to a linguistically motivated cogn…
ClusteringPOS induction with distributional and morphological information using a distance-dependent Chinese restaurant process
A Classification-Based Approach to Cognate Detection Combining Orthographic and Semantic Similarity Information
This paper presents proof-of-concept experiments for combining orthographic and semantic information to distinguish cognates from non-cognates. To this end, a context-independent gold standard is developed by manually la…
Binary ClassificationFormGeneral ClassificationSemantic Similarity+2Object proposal generation applying the distance dependent Chinese restaurant process
In application domains such as robotics, it is useful to represent the uncertainty related to the robot's belief about the state of its environment. Algorithms that only yield a single "best guess" as a result are not su…
Bayesian InferenceObjectObject DiscoveryObject Proposal GenerationTemporally-Reweighted Chinese Restaurant Process Mixtures for Clustering, Imputing, and Forecasting Multivariate Time Series
This article proposes a Bayesian nonparametric method for forecasting, imputation, and clustering in sparsely observed, multivariate time series data. The method is appropriate for jointly modeling hundreds of time serie…
ClusteringImputationTime SeriesTime Series Analysis