Understanding partition comparison indices based on counting object pairs
In unsupervised machine learning, agreement between partitions is commonly assessed with so-called external validity indices. Researchers tend to use and report indices that quantify agreement between two partitions for all clusters simultaneously. Commonly used examples are the Rand index and the adjusted Rand index. Since these overall measures give a general notion of what is going on, their values are usually hard to interpret. Three families of indices based on counting object pairs are analyzed. It is shown that the overall indices can be decomposed into indices that reflect the degree of agreement on the level of individual clusters. The overall indices based on the pair-counting approach are sensitive to cluster size imbalance: they tend to reflect the degree of agreement on the large clusters and provide little to no information on smaller clusters. Furthermore, the value of Rand-like indices is determined to a large extent by the number of pairs of objects that are not joined in either of the partitions.
Code (0)
등록된 구현이 없습니다.
Tasks
ObjectSimilar Papers 제목 키워드 기반
Adjusting for Chance Clustering Comparison Measures
Adjusted for chance measures are widely used to compare partitions/clusterings of the same data set. In particular, the Adjusted Rand Index (ARI) based on pair-counting, and the Adjusted Mutual Information (AMI) based on…
ClusteringUniformity in Heterogeneity:Diving Deep into Count Interval Partition for Crowd Counting
Recently, the problem of inaccurate learning targets in crowd counting draws increasing attention. Inspired by a few pioneering work, we solve this problem by trying to predict the indices of pre-defined interval bins of…
Crowd CountingQuantizationUniformity in Heterogeneity: Diving Deep Into Count Interval Partition for Crowd Counting
Recently, the problem of inaccurate learning targets in crowd counting draws increasing attention. Inspired by a few pioneering work, we solve this problem by trying to predict the indices of pre-defined interval bin…
Crowd CountingQuantizationOn the Use of Relative Validity Indices for Comparing Clustering Approaches
Relative Validity Indices (RVIs) such as the Silhouette Width Criterion and Davies Bouldin indices are the most widely used tools for evaluating and optimising clustering outcomes. Traditionally, their ability to rank co…
ClusteringDeep Clustering Evaluation: How to Validate Internal Clustering Validation Measures
Deep clustering, a method for partitioning complex, high-dimensional data using deep neural networks, presents unique evaluation challenges. Traditional clustering validation measures, designed for low-dimensional spaces…
ClusteringDeep Clustering