Unsupervised Decision Forest for Data Clustering and Density Estimation
An algorithm to improve performance parameter for unsupervised decision forest clustering and density estimation is presented. Specifically, a dual assignment parameter is introduced as a density estimator by combining Random Forest and Gaussian Mixture Model. The Random Forest method has been specifically applied to construct a robust affinity graph that provides information on the underlying structure of data objects used in clustering. The proposed algorithm differs from the commonly used spectral clustering methods where the computed distance metric is used to find similarities between data points. Experiments were conducted using five datasets. A comparison with six other state-of-the-art methods shows that our model is superior to existing approaches. Efficiency of the proposed model is in capturing the underlying structure for a given set of data points. The proposed method is also robust, and can discriminate between the complex features of data points among different clusters.
Code (0)
등록된 구현이 없습니다.
Tasks
ClusteringDensity EstimationMethods 이 논문이 사용한 방법론
Similar Papers 제목 키워드 기반
Density-based Clustering with Best-scored Random Forest
Single-level density-based approach has long been widely acknowledged to be a conceptually and mathematically convincing clustering method. In this paper, we propose an algorithm called "best-scored clustering forest" th…
ClusteringFederated unsupervised random forest for privacy-preserving patient stratification
In the realm of precision medicine, effective patient stratification and disease subtyping demand innovative methodologies tailored for multi-omics data. Clustering techniques applied to multi-omics data have become inst…
ClusteringFeature ImportancePrivacy PreservingUnsupervised anomaly detection algorithms on real-world data: how many do we need?
In this study we evaluate 32 unsupervised anomaly detection algorithms on 52 real-world multivariate tabular datasets, performing the largest comparison of unsupervised anomaly detection algorithms to date. On this colle…
Anomaly DetectionUnsupervised Anomaly DetectionUnsupervised detection of ash dieback disease (Hymenoscyphus fraxineus) using diffusion-based hyperspectral image clustering
Ash dieback (Hymenoscyphus fraxineus) is an introduced fungal disease that is causing the widespread death of ash trees across Europe. Remote sensing hyperspectral images encode rich structure that has been exploited for…
Clusteringhyperspectral image clusteringImage ClusteringImage Segmentation+1Feature graphs for interpretable unsupervised tree ensembles: centrality, interaction, and application in disease subtyping
Interpretable machine learning has emerged as central in leveraging artificial intelligence within high-stakes domains such as healthcare, where understanding the rationale behind model predictions is as critical as achi…
Clusteringfeature selectionInterpretable Machine Learning