paper-with-me

Papers

Robust Ultra-High-Dimensional Variable Selection With Correlated Structure Using Group Testing

2026-02-06 · Wanru Guo, Juan Xie, Binbin Wang, Weicong Chen, Xiaoyi Lu, Vipin Chaudhary, Curtis Tatsuoka arxiv

Background: High-dimensional genomic data exhibit strong group correlation structures that challenge conventional feature selection methods, which often assume feature independence or rely on pre-defined pathways and are sensitive to outliers and model misspecification. Methods: We propose the Dorfman screening framework, a multi-stage procedure that forms data-driven variable groups via hierarchical clustering, performs group and within-group hypothesis testing, and refines selection using elastic net or adaptive elastic net. Robust variants incorporate OGK-based covariance estimation, rank-based correlation, and Huber-weighted regression to handle contaminated and non-normal data. Results: In simulations, Dorfman-Sparse-Adaptive-EN performed best under normal conditions, while Robust-OGK-Dorfman-Adaptive-EN showed clear advantages under data contamination, outperforming classical Dorfman and competing methods. Applied to NSCLC gene expression data for trametinib response, robust Dorfman methods achieved the lowest prediction errors and enriched recovery of clinically relevant genes. Conclusions: The Dorfman framework provides an efficient and robust approach to genomic feature selection. Robust-OGK-Dorfman-Adaptive-EN offers strong performance under both ideal and contaminated conditions and scales to ultra-high-dimensional settings, making it well suited for modern genomic biomarker discovery.

📄 PDF Abstract BibTeX arXiv:2602.07258

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

A Two-Stage Variable Selection Approach for Correlated High Dimensional Predictors

2021-03-24 · Zhiyuan Li

When fitting statistical models, some predictors are often found to be correlated with each other, and functioning together. Many group variable selection methods are developed to select the groups of predictors that are…

ClusteringVariable SelectionVocal Bursts Intensity PredictionVocal Bursts Valence Prediction

Factor-Augmented Regularized Model for Hazard Regression

2022-10-03 · Pierre Bayle, Jianqing Fan

A prevalent feature of high-dimensional data is the dependence among covariates, and model selection is known to be challenging when covariates are highly correlated. To perform model selection for the high-dimensional C…

modelModel SelectionregressionSurvival Analysis

SPPCSO: Adaptive Penalized Estimation Method for High-Dimensional Correlated Data

2026-03-06 · Ying Hu, Hu Yang arxiv

With the rise of high-dimensional correlated data, multicollinearity poses a significant challenge to model stability, often leading to unstable estimation and reduced predictive accuracy. This work proposes the Single-P…

Efficient Clustering of Correlated Variables and Variable Selection in High-Dimensional Linear Models

2016-03-11 · Niharika Gauraha, Swapan K. Parui

In this paper, we introduce Adaptive Cluster Lasso(ACL) method for variable selection in high dimensional sparse regression models with strongly correlated variables. To handle correlated variables, the concept of cluste…

ClusteringVariable Selection

Error Controlled Feature Selection for Ultrahigh Dimensional and Highly Correlated Feature Space Using Deep Learning

2022-09-15 · Arkaprabha Ganguli, David Todem, Tapabrata Maiti

In recent years, deep learning has been at the center of analytics due to its impressive empirical success in analyzing complex data objects. Despite this success, most of the existing tools behave like black-box machine…

Deep Learningfeature selection