Two-Stage Robust and Sparse Distributed Statistical Inference for Large-Scale Data
In this paper, we address the problem of conducting statistical inference in settings involving large-scale data that may be high-dimensional and contaminated by outliers. The high volume and dimensionality of the data require distributed processing and storage solutions. We propose a two-stage distributed and robust statistical inference procedures coping with high-dimensional models by promoting sparsity. In the first stage, known as model selection, relevant predictors are locally selected by applying robust Lasso estimators to the distinct subsets of data. The variable selections from each computation node are then fused by a voting scheme to find the sparse basis for the complete data set. It identifies the relevant variables in a robust manner. In the second stage, the developed statistically robust and computationally efficient bootstrap methods are employed. The actual inference constructs confidence intervals, finds parameter estimates and quantifies standard deviation. Similar to stage 1, the results of local inference are communicated to the fusion center and combined there. By using analytical methods, we establish the favorable statistical properties of the robust and computationally efficient bootstrap methods including consistency for a fixed number of predictors, and robustness. The proposed two-stage robust and distributed inference procedures demonstrate reliable performance and robustness in variable selection, finding confidence intervals and bootstrap approximations of standard deviations even when data is high-dimensional and contaminated by outliers.
Code (0)
등록된 구현이 없습니다.
Tasks
Model SelectionVariable SelectionSimilar Papers 제목 키워드 기반
Minimax and Communication-Efficient Distributed Best Subset Selection with Oracle Property
The explosion of large-scale data in fields such as finance, e-commerce, and social media has outstripped the processing capabilities of single-machine systems, driving the need for distributed statistical inference meth…
Distributed Semi-Supervised Sparse Statistical Inference
The debiased estimator is a crucial tool in statistical inference for high-dimensional model parameters. However, constructing such an estimator involves estimating the high-dimensional inverse Hessian matrix, incurring …
DDAC-SpAM: A Distributed Algorithm for Fitting High-dimensional Sparse Additive Models with Feature Division and Decorrelation
Distributed statistical learning has become a popular technique for large-scale data analysis. Most existing work in this area focuses on dividing the observations, but we propose a new algorithm, DDAC-SpAM, which divide…
Additive modelsfeature selectionVocal Bursts Intensity PredictionFocal and Connectomic Mapping of Transiently Disrupted Brain Function
The distributed nature of the neural substrate, and the difficulty of establishing necessity from correlative data, combine to render the mapping of brain function a far harder task than it seems. Methods capable of comb…
Statistically Guided Divide-and-Conquer for Sparse Factorization of Large Matrix
The sparse factorization of a large matrix is fundamental in modern statistical learning. In particular, the sparse singular value decomposition and its variants have been utilized in multivariate regression, factor anal…
Computational EfficiencyregressionTime Series Analysis