paper-with-me

홈 › Papers

Visualizing the Finer Cluster Structure of Large-Scale and High-Dimensional Data

2020-07-17 · Yu Liang, Arin Chaudhuri, Haoyu Wang

Dimension reduction and visualization of high-dimensional data have become very important research topics because of the rapid growth of large databases in data science. In this paper, we propose using a generalized sigmoid function to model the distance similarity in both high- and low-dimensional spaces. In particular, the parameter b is introduced to the generalized sigmoid function in low-dimensional space, so that we can adjust the heaviness of the function tail by changing the value of b. Using both simulated and real-world data sets, we show that our proposed method can generate visualization results comparable to those of uniform manifold approximation and projection (UMAP), which is a newly developed manifold learning technique with fast running speed, better global structure, and scalability to massive data sets. In addition, according to the purpose of the study and the data structure, we can decrease or increase the value of b to either reveal the finer cluster structure of the data or maintain the neighborhood continuity of the embedding for better visualization. Finally, we use domain knowledge to demonstrate that the finer subclusters revealed with small values of b are meaningful.

📄 PDF Abstract BibTeX arXiv:2007.08711

Code (0)

등록된 구현이 없습니다.

Tasks

Dimensionality Reduction

Similar Papers 제목 키워드 기반

TiVy: Time Series Visual Summary for Scalable Visualization

2025-07-25 · Gromit Yeuk-Yin Chan, Luis Gustavo Nonato, Themis Palpanas, Cláudio T. Silva 외 arxiv

Visualizing multiple time series presents fundamental tradeoffs between scalability and visual clarity. Time series capture the behavior of many large-scale real-world processes, from stock market trends to urban activit…

Visualizing Overlapping Biclusterings and Boolean Matrix Factorizations

2023-07-14 · Thibault Marette, Pauli Miettinen, Stefan Neumann

Finding (bi-)clusters in bipartite graphs is a popular data analysis approach. Analysts typically want to visualize the clusters, which is simple as long as the clusters are disjoint. However, many modern algorithms find…

Visualizing Classification Structure of Large-Scale Classifiers

2020-07-12 · Bilal Alsallakh, Zhixin Yan, Shabnam Ghaffarzadegan, Zeng Dai 외

We propose a measure to compute class similarity in large-scale classification based on prediction scores. Such measure has not been formally pro-posed in the literature. We show how visualizing the class similarity matr…

ClassificationGeneral Classification

Heavy-tailed kernels reveal a finer cluster structure in t-SNE visualisations

2019-02-15 · Dmitry Kobak, George Linderman, Stefan Steinerberger, Yuval Kluger 외

T-distributed stochastic neighbour embedding (t-SNE) is a widely used data visualisation technique. It differs from its predecessor SNE by the low-dimensional similarity kernel: the Gaussian kernel was replaced by the he…

XFlowMap: Cross-Scale Generalization and Mapping of Massive Origin-Destination Data

2026-04-23 · Diansheng Guo, Hai Jin arxiv

Mapping large origin-destination (OD) datasets remains challenging because flow maps become cluttered, meaningful patterns occur at multiple spatial scales, and existing flow-mapping approaches frequently rely on predefi…