Unsupervised Feature Selection for the k-means Clustering Problem
We present a novel feature selection algorithm for the $k$-means clustering problem. Our algorithm is randomized and, assuming an accuracy parameter $\epsilon \in (0,1)$, selects and appropriately rescales in an unsupervised manner $\Theta(k \log(k / \epsilon) / \epsilon^2)$ features from a dataset of arbitrary dimensions. We prove that, if we run any $\gamma$-approximate $k$-means algorithm ($\gamma \geq 1$) on the features selected using our method, we can find a $(1+(1+\epsilon)\gamma)$-approximate partition with high probability.
Code (0)
등록된 구현이 없습니다.
Tasks
Clusteringfeature selectionSimilar Papers 제목 키워드 기반
K-means Derived Unsupervised Feature Selection using Improved ADMM
Feature selection is important for high-dimensional data analysis and is non-trivial in unsupervised learning problems such as dimensionality reduction and clustering. The goal of unsupervised feature selection is findin…
ClusteringDimensionality Reductionfeature selectionFeature selection or extraction decision process for clustering using PCA and FRSD
This paper concerns the critical decision process of extracting or selecting the features before applying a clustering algorithm. It is not obvious to evaluate the importance of the features since the most popular method…
ClusteringDimensionality Reductionfeature selectioni-IF-Learn: Iterative Feature Selection and Unsupervised Learning for High-Dimensional Complex Data
Unsupervised learning of high-dimensional data is challenging due to irrelevant or noisy features obscuring underlying structures. It's common that only a few features, called the influential features, meaningfully defin…
Deep ClusteringClustering with feature selection using alternating minimization, Application to computational biology
This paper deals with unsupervised clustering with feature selection. The problem is to estimate both labels and a sparse projection matrix of weights. To address this combinatorial non-convex problem maintaining a stric…
Clusteringfeature selectionUnsupervised Gene Expression Data using Enhanced Clustering Method
Microarrays are made it possible to simultaneously monitor the expression profiles of thousands of genes under various experimental conditions. Identification of co-expressed genes and coherent patterns is the central go…
Clusteringfeature selection