paper-with-me

Papers

Distributed Bootstrap for Simultaneous Inference Under High Dimensionality

2021-02-19 · Yang Yu, Shih-Kang Chao, Guang Cheng

We propose a distributed bootstrap method for simultaneous inference on high-dimensional massive data that are stored and processed with many machines. The method produces an $\ell_\infty$-norm confidence region based on a communication-efficient de-biased lasso, and we propose an efficient cross-validation approach to tune the method at every iteration. We theoretically prove a lower bound on the number of communication rounds $\tau_{\min}$ that warrants the statistical accuracy and efficiency. Furthermore, $\tau_{\min}$ only increases logarithmically with the number of workers and the intrinsic dimensionality, while nearly invariant to the nominal dimensionality. We test our theory by extensive simulation studies, and a variable screening task on a semi-synthetic dataset based on the US Airline On-Time Performance dataset. The code to reproduce the numerical results is available at GitHub: https://github.com/skchao74/Distributed-bootstrap.

📄 PDF Abstract BibTeX arXiv:2102.10080

Code (1)

skchao74/Distributed-bootstrap 공식 구현

Tasks

Vocal Bursts Intensity Prediction

Similar Papers 제목 키워드 기반

Simultaneous Inference for Massive Data: Distributed Bootstrap

2020-02-19 · ICML 2020 1 · Yang Yu, Shih-Kang Chao, Guang Cheng

In this paper, we propose a bootstrap method applied to massive data processed distributedly in a large number of machines. This new method is computationally efficient in that we bootstrap on the master machine without …

Comments on `High-dimensional simultaneous inference with the bootstrap'

2017-05-06 · Jelena Bradic, Yinchu Zhu

We provide comments on the article "High-dimensional simultaneous inference with the bootstrap" by Ruben Dezeure, Peter Buhlmann and Cun-Hui Zhang.

Vocal Bursts Intensity Prediction

Two-Stage Robust and Sparse Distributed Statistical Inference for Large-Scale Data

2022-08-17 · Emadaldin Mozafari-Majd, Visa Koivunen

In this paper, we address the problem of conducting statistical inference in settings involving large-scale data that may be high-dimensional and contaminated by outliers. The high volume and dimensionality of the data r…

Model SelectionVariable Selection

Iterative Distributed Multinomial Regression

2024-12-02 · Yanqin Fan, Yigit Okar, Xuetao Shi

This article introduces an iterative distributed computing estimator for the multinomial logistic regression model with large choice sets. Compared to the maximum likelihood estimator, the proposed iterative distributed …

Computational EfficiencyDistributed Computingregression

Robust, scalable and fast bootstrap method for analyzing large scale data

2015-04-09 · Shahab Basiri, Esa Ollila, Visa Koivunen

In this paper we address the problem of performing statistical inference for large scale data sets i.e., Big Data. The volume and dimensionality of the data may be so high that it cannot be processed or stored in a singl…