Heavy-tailed kernels reveal a finer cluster structure in t-SNE visualisations
T-distributed stochastic neighbour embedding (t-SNE) is a widely used data visualisation technique. It differs from its predecessor SNE by the low-dimensional similarity kernel: the Gaussian kernel was replaced by the heavy-tailed Cauchy kernel, solving the "crowding problem" of SNE. Here, we develop an efficient implementation of t-SNE for a $t$-distribution kernel with an arbitrary degree of freedom $\nu$, with $\nu\to\infty$ corresponding to SNE and $\nu=1$ corresponding to the standard t-SNE. Using theoretical analysis and toy examples, we show that $\nu<1$ can further reduce the crowding problem and reveal finer cluster structure that is invisible in standard t-SNE. We further demonstrate the striking effect of heavier-tailed kernels on large real-life data sets such as MNIST, single-cell RNA-sequencing data, and the HathiTrust library. We use domain knowledge to confirm that the revealed clusters are meaningful. Overall, we argue that modifying the tail heaviness of the t-SNE kernel can yield additional insight into the cluster structure of the data.
Code (2)
Similar Papers 제목 키워드 기반
Cluster Analysis with Resampling for Validation and Exploration (CARVE)
Clustering is widely used across the sciences as the foundation for downstream data-driven scientific discoveries. However, clustering results are highly sensitive to the choice of algorithm, preprocessing, and the numbe…
Convergence of Heavy-Tailed Hawkes Processes and the Microstructure of Rough Volatility
We establish the weak convergence of the intensity of a nearly-unstable Hawkes process with heavy-tailed kernel. Our result is used to derive a scaling limit for a financial market model where orders to buy or sell an as…
Robust convex biclustering with a tuning-free method
Biclustering is widely used in different kinds of fields including gene information analysis, text mining, and recommendation system by effectively discovering the local correlation between samples and features. However,…
MIM-Refiner: A Contrastive Learning Boost from Intermediate Pre-Trained Representations
We introduce MIM (Masked Image Modeling)-Refiner, a contrastive learning boost for pre-trained MIM models. MIM-Refiner is motivated by the insight that strong representations within MIM models generally reside in interme…
Contrastive LearningImage ClusteringSelf-Supervised Image ClassificationSemantic SegmentationReal Elliptically Skewed Distributions and Their Application to Robust Cluster Analysis
This article proposes a new class of Real Elliptically Skewed (RESK) distributions and associated clustering algorithms that allow for integrating robustness and skewness into a single unified cluster analysis framework.…
Clustering