paper-with-me

홈 › Papers

Out-of-Core Dimensionality Reduction for Large Data via Out-of-Sample Extensions

2024-08-07 · Luca Reichmann, David Hägele, Daniel Weiskopf

Dimensionality reduction (DR) is a well-established approach for the visualization of high-dimensional data sets. While DR methods are often applied to typical DR benchmark data sets in the literature, they might suffer from high runtime complexity and memory requirements, making them unsuitable for large data visualization especially in environments outside of high-performance computing. To perform DR on large data sets, we propose the use of out-of-sample extensions. Such extensions allow inserting new data into existing projections, which we leverage to iteratively project data into a reference projection that consists only of a small manageable subset. This process makes it possible to perform DR out-of-core on large data, which would otherwise not be possible due to memory and runtime limitations. For metric multidimensional scaling (MDS), we contribute an implementation with out-of-sample projection capability since typical software libraries do not support it. We provide an evaluation of the projection quality of five common DR algorithms (MDS, PCA, t-SNE, UMAP, and autoencoders) using quality metrics from the literature and analyze the trade-off between the size of the reference set and projection quality. The runtime behavior of the algorithms is also quantified with respect to reference set size, out-of-sample batch size, and dimensionality of the data sets. Furthermore, we compare the out-of-sample approach to other recently introduced DR methods, such as PaCMAP and TriMAP, which claim to handle larger data sets than traditional approaches. To showcase the usefulness of DR on this large scale, we contribute a use case where we analyze ensembles of streamlines amounting to one billion projected instances.

📄 PDF Abstract BibTeX arXiv:2408.04129

Code (0)

등록된 구현이 없습니다.

Tasks

Data VisualizationDimensionality Reduction

Methods 이 논문이 사용한 방법론

SET Dynamic Sparse Training method where weight mask is updated randomly periodically
PCA Principle Components Analysis (PCA) is an unsupervised method primary used for dimensionality reduction within machine learning. PCA is calculated via a singular value…

Similar Papers 제목 키워드 기반

Projecting "better than randomly": How to reduce the dimensionality of very large datasets in a way that outperforms random projections

2019-01-03 · Michael Wojnowicz, Di Zhang, Glenn Chisholm, Xuan Zhao 외

For very large datasets, random projections (RP) have become the tool of choice for dimensionality reduction. This is due to the computational complexity of principal component analysis. However, the recent development o…

Dimensionality ReductionGeneral ClassificationMalware Classification

PCA-Based Out-of-Sample Extension for Dimensionality Reduction

2015-11-03 · Yariv Aizenbud, Amit Bermanis, Amir Averbuch

Dimensionality reduction methods are very common in the field of high dimensional data analysis. Typically, algorithms for dimensionality reduction are computationally expensive. Therefore, their applications for the ana…

Dimensionality Reduction

Joint Dimensionality Reduction for Separable Embedding Estimation

2021-01-14 · Yanjun Li, Bihan Wen, Hao Cheng, Yoram Bresler

Low-dimensional embeddings for data from disparate sources play critical roles in multi-modal machine learning, multimedia information retrieval, and bioinformatics. In this paper, we propose a supervised dimensionality …

Dimensionality Reductionfeature selectionInformation Retrievalregression+2

Feature Dimensionality Reduction for Video Affect Classification: A Comparative Study

2018-08-08 · Chenfeng Guo, Dongrui Wu

Affective computing has become a very important research area in human-machine interaction. However, affects are subjective, subtle, and uncertain. So, it is very difficult to obtain a large number of labeled training sa…

ClassificationDimensionality ReductionGeneral Classification

CCP: Correlated Clustering and Projection for Dimensionality Reduction

2022-06-08 · Yuta Hozumi, Rui Wang, Guo-Wei Wei

Most dimensionality reduction methods employ frequency domain representations obtained from matrix diagonalization and may not be efficient for large datasets with relatively high intrinsic dimensions. To address this ch…

ClusteringDimensionality Reduction