paper-with-me

Papers

Optimal Distributed Subsampling for Maximum Quasi-Likelihood Estimators with Massive Data

2020-05-21 · Jun Yu, HaiYing Wang, Mingyao Ai, Huiming Zhang

Nonuniform subsampling methods are effective to reduce computational burden and maintain estimation efficiency for massive data. Existing methods mostly focus on subsampling with replacement due to its high computational efficiency. If the data volume is so large that nonuniform subsampling probabilities cannot be calculated all at once, then subsampling with replacement is infeasible to implement. This paper solves this problem using Poisson subsampling. We first derive optimal Poisson subsampling probabilities in the context of quasi-likelihood estimation under the A- and L-optimality criteria. For a practically implementable algorithm with approximated optimal subsampling probabilities, we establish the consistency and asymptotic normality of the resultant estimators. To deal with the situation that the full data are stored in different blocks or at multiple locations, we develop a distributed subsampling framework, in which statistics are computed simultaneously on smaller partitions of the full data. Asymptotic properties of the resultant aggregated estimator are investigated. We illustrate and evaluate the proposed strategies through numerical experiments on simulated and real data sets.

📄 PDF Abstract BibTeX arXiv:2005.10435

Code (0)

등록된 구현이 없습니다.

Tasks

Computational Efficiency

Similar Papers 제목 키워드 기반

Maximum sampled conditional likelihood for informative subsampling

2020-11-11 · Haiying Wang, Jae Kwang Kim

Subsampling is a computationally effective approach to extract information from massive data sets when computing resources are limited. After a subsample is taken from the full data, most available methods use an inverse…

Optimal Subsampling for Large Sample Logistic Regression

2017-02-03 · HaiYing Wang, Rong Zhu, Ping Ma

For massive data, the family of subsampling algorithms is popular to downsize the data volume and reduce computational burden. Existing studies focus on approximating the ordinary least squares estimate in linear regress…

regression

Asymptotic equivalence of Principal Components and Quasi Maximum Likelihood estimators in Large Approximate Factor Models

2023-07-19 · Matteo Barigozzi

This paper investigates the properties of Quasi Maximum Likelihood estimation of an approximate factor model for an $n$-dimensional vector of stationary time series. We prove that the factor loadings estimated by Quasi M…

regressionTime SeriesTime Series Regression

Communication-Efficient Distributed Statistical Inference

2016-05-25 · Michael. I. Jordan, Jason D. Lee, Yun Yang

We present a Communication-efficient Surrogate Likelihood (CSL) framework for solving distributed statistical inference problems. CSL provides a communication-efficient surrogate to the global likelihood that can be used…

Bayesian InferenceComputational Efficiency

The block-Poisson estimator for optimally tuned exact subsampling MCMC

2016-03-27 · Matias Quiroz, Minh-Ngoc Tran, Mattias Villani, Robert Kohn 외

Speeding up Markov Chain Monte Carlo (MCMC) for datasets with many observations by data subsampling has recently received considerable attention. A pseudo-marginal MCMC method is proposed that estimates the likelihood by…