Simultaneous Inference for Massive Data: Distributed Bootstrap
In this paper, we propose a bootstrap method applied to massive data processed distributedly in a large number of machines. This new method is computationally efficient in that we bootstrap on the master machine without over-resampling, typically required by existing methods \cite{kleiner2014scalable,sengupta2016subsampled}, while provably achieving optimal statistical efficiency with minimal communication. Our method does not require repeatedly re-fitting the model but only applies multiplier bootstrap in the master machine on the gradients received from the worker machines. Simulations validate our theory.
Code (0)
등록된 구현이 없습니다.
Similar Papers 제목 키워드 기반
Distributed Bootstrap for Simultaneous Inference Under High Dimensionality
We propose a distributed bootstrap method for simultaneous inference on high-dimensional massive data that are stored and processed with many machines. The method produces an $\ell_\infty$-norm confidence region based on…
Vocal Bursts Intensity PredictionComments on `High-dimensional simultaneous inference with the bootstrap'
We provide comments on the article "High-dimensional simultaneous inference with the bootstrap" by Ruben Dezeure, Peter Buhlmann and Cun-Hui Zhang.
Vocal Bursts Intensity PredictionStatistical inference in massive datasets by empirical likelihood
In this paper, we propose a new statistical inference method for massive data sets, which is very simple and efficient by combining divide-and-conquer method and empirical likelihood. Compared with two popular methods (t…
Gaussian Differential Private Bootstrap by Subsampling
Bootstrap is a common tool for quantifying uncertainty in data analysis. However, besides additional computational costs in the application of the bootstrap on massive data, a challenging problem in bootstrap based infer…
Uncertainty QuantificationTwo-Stage Robust and Sparse Distributed Statistical Inference for Large-Scale Data
In this paper, we address the problem of conducting statistical inference in settings involving large-scale data that may be high-dimensional and contaminated by outliers. The high volume and dimensionality of the data r…
Model SelectionVariable Selection