paper-with-me

홈 › Papers

Kernel k-Groups via Hartigan's Method

2017-10-26 · Guilherme França, Maria L. Rizzo, Joshua T. Vogelstein

Energy statistics was proposed by Sz\' ekely in the 80's inspired by Newton's gravitational potential in classical mechanics and it provides a model-free hypothesis test for equality of distributions. In its original form, energy statistics was formulated in Euclidean spaces. More recently, it was generalized to metric spaces of negative type. In this paper, we consider a formulation for the clustering problem using a weighted version of energy statistics in spaces of negative type. We show that this approach leads to a quadratically constrained quadratic program in the associated kernel space, establishing connections with graph partitioning problems and kernel methods in machine learning. To find local solutions of such an optimization problem, we propose kernel k-groups, which is an extension of Hartigan's method to kernel spaces. Kernel k-groups is cheaper than spectral clustering and has the same computational cost as kernel k-means (which is based on Lloyd's heuristic) but our numerical results show an improved performance, especially in higher dimensions. Moreover, we verify the efficiency of kernel k-groups in community detection in sparse stochastic block models which has fascinating applications in several areas of science.

📄 PDF Abstract BibTeX arXiv:1710.09859

Code (0)

등록된 구현이 없습니다.

Tasks

ClusteringCommunity Detectiongraph partitioning

Methods 이 논문이 사용한 방법론

Spectral Clustering Spectral clustering has attracted increasing attention due to the promising ability in dealing with nonlinearly separable datasets [15], [16]. In spectral clustering, the…

Similar Papers 제목 키워드 기반

An effective variant of the Hartigan $k$-means algorithm

2026-04-23 · François Clément, Stefan Steinerberger arxiv

The k-means problem is perhaps the classical clustering problem and often synonymous with Lloyd's algorithm (1957). It has become clear that Hartigan's algorithm (1975) gives better results in almost all cases, Telgarsky…

Beyond Hartigan Consistency: Merge Distortion Metric for Hierarchical Clustering

2015-06-21 · Justin Eldridge, Mikhail Belkin, Yusu Wang

Hierarchical clustering is a popular method for analyzing data which associates a tree to a dataset. Hartigan consistency has been used extensively as a framework to analyze such clustering algorithms from a statistical …

Clustering

The Catastrophic Failure of The k-Means Algorithm in High Dimensions, and How Hartigan's Algorithm Avoids It

2026-02-10 · Roy R. Lederman, David Silva-Sánchez, Ziling Chen, Gilles Mordant 외 arxiv

Lloyd's k-means algorithm is one of the most widely used clustering methods. We prove that in high-dimensional, high-noise settings, the algorithm exhibits catastrophic failure: with high probability, essentially every p…

Further heuristics for $k$-means: The merge-and-split heuristic and the $(k,l)$-means

2014-06-23 · Frank Nielsen, Richard Nock

Finding the optimal $k$-means clustering is NP-hard in general and many heuristics have been designed for minimizing monotonically the $k$-means objective. We first show how to extend Lloyd's batched relocation heuristic…

Clustering

Moving Up the Cluster Tree with the Gradient Flow

2021-09-17 · Ery Arias-Castro, Wanli Qiao

The paper establishes a strong correspondence between two important clustering approaches that emerged in the 1970's: clustering by level sets or cluster tree as proposed by Hartigan and clustering by gradient lines or g…

Clustering