paper-with-me

Papers

Clustering categorical data via ensembling dissimilarity matrices

2015-06-26 · Saeid Amiri, Bertrand Clarke, Jennifer Clarke

We present a technique for clustering categorical data by generating many dissimilarity matrices and averaging over them. We begin by demonstrating our technique on low dimensional categorical data and comparing it to several other techniques that have been proposed. Then we give conditions under which our method should yield good results in general. Our method extends to high dimensional categorical data of equal lengths by ensembling over many choices of explanatory variables. In this context we compare our method with two other methods. Finally, we extend our method to high dimensional categorical data vectors of unequal length by using alignment techniques to equalize the lengths. We give examples to show that our method continues to provide good results, in particular, better in the context of genome sequences than clusterings suggested by phylogenetic trees.

📄 PDF Abstract BibTeX arXiv:1506.07930

Code (0)

등록된 구현이 없습니다.

Tasks

Clustering

Similar Papers 제목 키워드 기반

Mixed-type Distance Shrinkage and Selection for Clustering via Kernel Metric Learning

2023-06-02 · Jesse S. Ghashti, John R. J. Thompson

Distance-based clustering and classification are widely used in various fields to group mixed numeric and categorical data. In many algorithms, a predefined distance measurement is used to cluster data points based on th…

AttributeClusteringMetric Learning

Learnable Weighting of Intra-Attribute Distances for Categorical Data Clustering with Nominal and Ordinal Attributes

2026-07-06 · Yiqun Zhang, Yiu-ming Cheung arxiv

The success of categorical data clustering generally much relies on the distance metric that measures the dissimilarity degree between two objects. However, most of the existing clustering methods treat the two categoric…

Clustering in Hilbert simplex geometry

2017-04-03 · Frank Nielsen, Ke Sun

Clustering categorical distributions in the finite-dimensional probability simplex is a fundamental task met in many applications dealing with normalized histograms. Traditionally, the differential-geometric structures o…

Clustering

Impact of Event Encoding and Dissimilarity Measures on Traffic Crash Characterization Based on Sequence of Events

2023-02-22 · Yu Song, Madhav V. Chitturi, David A. Noyce

Crash sequence analysis has been shown in prior studies to be useful for characterizing crashes and identifying safety countermeasures. Sequence analysis is highly domain-specific, but its various techniques have not bee…

Clustering

Estimating the Optimal Number of Clusters in Categorical Data Clustering by Silhouette Coefficient

2025-01-26 · Duy-Tai Dinh, Tsutomu Fujinami, Van-Nam Huynh

The problem of estimating the number of clusters (say k) is one of the major challenges for the partitional clustering. This paper proposes an algorithm named k-SCC to estimate the optimal k in categorical data clusterin…

Categorical data clusteringClusteringDensity Estimation