Ridge partial correlation screening for ultrahigh-dimensional data
Variable selection in ultrahigh-dimensional linear regression is challenging due to its high computational cost. Therefore, a screening step is usually conducted before variable selection to significantly reduce the dimension. Here we propose a novel and simple screening method based on ordering the absolute sample ridge partial correlations. The proposed method takes into account not only the ridge regularized estimates of the regression coefficients but also the ridge regularized partial variances of the predictor variables providing sure screening property without strong assumptions on the marginal correlations. Simulation study and a real data analysis show that the proposed method has a competitive performance compared with the existing screening procedures. A publicly available software implementing the proposed screening accompanies the article.
Code (0)
등록된 구현이 없습니다.
Tasks
regressionVariable SelectionMethods 이 논문이 사용한 방법론
Similar Papers 제목 키워드 기반
ExSIS: Extended Sure Independence Screening for Ultrahigh-dimensional Linear Models
Statistical inference can be computationally prohibitive in ultrahigh-dimensional linear models. Correlation-based variable screening, in which one leverages marginal correlations for removal of irrelevant variables from…
Feature space reduction method for ultrahigh-dimensional, multiclass data: Random forest-based multiround screening (RFMS)
In recent years, numerous screening methods have been published for ultrahigh-dimensional data that contain hundreds of thousands of features; however, most of these features cannot handle data with thousands of classes.…
Classification with Ultrahigh-Dimensional Features
Although much progress has been made in classification with high-dimensional features \citep{Fan_Fan:2008, JGuo:2010, CaiSun:2014, PRXu:2014}, classification with ultrahigh-dimensional features, wherein the features much…
ClassificationGeneral ClassificationCovariance-Insured Screening
Modern bio-technologies have produced a vast amount of high-throughput data with the number of predictors far greater than the sample size. In order to identify more novel biomarkers and understand biological mechanisms,…
An adaptive subsampling method for large-sample feature screening
We consider the sure independence screening (SIS) method, a standard feature screening approach that aims to eliminate non-informative features in ultrahigh-dimensional datasets. Although effective, SIS incurs a computat…
Computational Efficiency