paper-with-me

홈 › Papers

Classes are not Clusters: Improving Label-based Evaluation of Dimensionality Reduction

2023-08-01 · Hyeon Jeon, Yun-Hsin Kuo, Michaël Aupetit, Kwan-Liu Ma, Jinwook Seo

A common way to evaluate the reliability of dimensionality reduction (DR) embeddings is to quantify how well labeled classes form compact, mutually separated clusters in the embeddings. This approach is based on the assumption that the classes stay as clear clusters in the original high-dimensional space. However, in reality, this assumption can be violated; a single class can be fragmented into multiple separated clusters, and multiple classes can be merged into a single cluster. We thus cannot always assure the credibility of the evaluation using class labels. In this paper, we introduce two novel quality measures -- Label-Trustworthiness and Label-Continuity (Label-T&C) -- advancing the process of DR evaluation based on class labels. Instead of assuming that classes are well-clustered in the original space, Label-T&C work by (1) estimating the extent to which classes form clusters in the original and embedded spaces and (2) evaluating the difference between the two. A quantitative evaluation showed that Label-T&C outperform widely used DR evaluation measures (e.g., Trustworthiness and Continuity, Kullback-Leibler divergence) in terms of the accuracy in assessing how well DR embeddings preserve the cluster structure, and are also scalable. Moreover, we present case studies demonstrating that Label-T&C can be successfully used for revealing the intrinsic characteristics of DR techniques and their hyperparameters.

📄 PDF Abstract BibTeX arXiv:2308.00278

Code (1)

hj-n/ltnc 공식 구현

Tasks

Dimensionality Reduction

Similar Papers 제목 키워드 기반

Supervised Stochastic Neighbor Embedding Using Contrastive Learning

2023-09-15 · Yi Zhang

Stochastic neighbor embedding (SNE) methods $t$-SNE, UMAP are two most popular dimensionality reduction methods for data visualization. Contrastive learning, especially self-supervised contrastive learning (SSCL), has sh…

Contrastive LearningData VisualizationDimensionality Reduction

Deep Autoencoders for Dimensionality Reduction of High-Content Screening Data

2015-01-07 · Lee Zamparo, Zhaolei Zhang

High-content screening uses large collections of unlabeled cell image data to reason about genetics or cell biology. Two important tasks are to identify those cells which bear interesting phenotypes, and to identify sub-…

ClusteringDimensionality ReductionVocal Bursts Intensity Prediction

A Fuzzy Approach for Feature Evaluation and Dimensionality Reduction to Improve the Quality of Web Usage Mining Results

2015-09-01 · Zahid Ansari, M. F. Azeem, A. Vinaya Babu, Waseem Ahmed

Web Usage Mining is the application of data mining techniques to web usage log repositories in order to discover the usage patterns that can be used to analyze the users navigational behavior. During the preprocessing st…

ClusteringDimensionality Reductionfeature selection

Assessing the impact of dimensionality reduction on clustering performance -- a systematic study

2026-04-23 · Ousmane Assani-Amate, Mohammadreza Bakhtyari, Émilie Roy, Vladimir Makarenkov arxiv

Dimensionality reduction is a critical preprocessing step for clustering high-dimensional data, yet comprehensive evaluation of its impact across diverse methods and data types remains limited. In this study, we systemat…

Dimensionality Reduction

IT-map: an Effective Nonlinear Dimensionality Reduction Method for Interactive Clustering

2015-01-26 · Teng Qiu, Yong-Jie Li

Scientists in many fields have the common and basic need of dimensionality reduction: visualizing the underlying structure of the massive multivariate data in a low-dimensional space. However, many dimensionality reducti…

ClusteringDimensionality Reduction