paper-with-me

홈 › Papers

A Unified Framework for Variable Selection in Model-Based Clustering with Missing Not at Random

2025-05-25 · Binh H. Ho, Long Nguyen Chi, TrungTin Nguyen, Binh T. Nguyen, Van Ha Hoang, Christopher Drovandi

Model-based clustering integrated with variable selection is a powerful tool for uncovering latent structures within complex data. However, its effectiveness is often hindered by challenges such as identifying relevant variables that define heterogeneous subgroups and handling data that are missing not at random, a prevalent issue in fields like transcriptomics. While several notable methods have been proposed to address these problems, they typically tackle each issue in isolation, thereby limiting their flexibility and adaptability. This paper introduces a unified framework designed to address these challenges simultaneously. Our approach incorporates a data-driven penalty matrix into penalized clustering to enable more flexible variable selection, along with a mechanism that explicitly models the relationship between missingness and latent class membership. We demonstrate that, under certain regularity conditions, the proposed framework achieves both asymptotic consistency and selection consistency, even in the presence of missing data. This unified strategy significantly enhances the capability and efficiency of model-based clustering, advancing methodologies for identifying informative variables that define homogeneous subgroups in the presence of complex missing data patterns. The performance of the framework, including its computational efficiency, is evaluated through simulations and demonstrated using both synthetic and real-world transcriptomic datasets.

📄 PDF Abstract BibTeX arXiv:2505.19093

Code (0)

등록된 구현이 없습니다.

Tasks

ClusteringComputational EfficiencyVariable Selection

Similar Papers 제목 키워드 기반

TRUST-FS: Tensorized Reliable Unsupervised Multi-View Feature Selection for Incomplete Data

2025-09-16 · Minghui Lu, Yanyong Huang, Minbo Ma, Jinyuan Chang 외 arxiv

Multi-view unsupervised feature selection (MUFS), which selects informative features from multi-view unlabeled data, has attracted increasing research interest in recent years. Although great efforts have been devoted to…

MIBoost: A gradient boosting algorithm for variable selection after multiple imputation

2025-07-29 · Robert Kuchen arxiv

Statistical learning methods for automated variable selection, such as the Least Absolute Shrinkage and Selection Operator (LASSO), elastic nets, and gradient boosting, have become increasingly popular tools for building…

URRL-IMVC: Unified and Robust Representation Learning for Incomplete Multi-View Clustering

2024-07-12 · Ge Teng, Ting Mao, Chen Shen, Xiang Tian 외

Incomplete multi-view clustering (IMVC) aims to cluster multi-view data that are only partially available. This poses two main challenges: effectively leveraging multi-view information and mitigating the impact of missin…

ClusteringContrastive LearningData AugmentationImputation+2

Missing Value Knockoffs

2022-02-26 · Deniz Koyuncu, Bülent Yener

One limitation of the most statistical/machine learning-based variable selection approaches is their inability to control the false selections. A recently introduced framework, model-x knockoffs, provides that to a wide …

ImputationMissing ValuesVariable Selection

Variable selection for clustering with Gaussian mixture models: state of the art

2017-01-31 · Abdelghafour Talibi, Boujemâa Achchab, Rafik Lasri

The mixture models have become widely used in clustering, given its probabilistic framework in which its based, however, for modern databases that are characterized by their large size, these models behave disappointingl…

ClusteringVariable Selection