paper-with-me

Papers

Data-adaptive statistics for multiple hypothesis testing in high-dimensional settings

2017-04-24 · Weixin Cai, Nima S. Hejazi, Alan E. Hubbard

Current statistical inference problems in areas like astronomy, genomics, and marketing routinely involve the simultaneous testing of thousands -- even millions -- of null hypotheses. For high-dimensional multivariate distributions, these hypotheses may concern a wide range of parameters, with complex and unknown dependence structures among variables. In analyzing such hypothesis testing procedures, gains in efficiency and power can be achieved by performing variable reduction on the set of hypotheses prior to testing. We present in this paper an approach using data-adaptive multiple testing that serves exactly this purpose. This approach applies data mining techniques to screen the full set of covariates on equally sized partitions of the whole sample via cross-validation. This generalized screening procedure is used to create average ranks for covariates, which are then used to generate a reduced (sub)set of hypotheses, from which we compute test statistics that are subsequently subjected to standard multiple testing corrections. The principal advantage of this methodology lies in its providing valid statistical inference without the \textit{a priori} specifying which hypotheses will be tested. Here, we present the theoretical details of this approach, confirm its validity via a simulation study, and exemplify its use by applying it to the analysis of data on microRNA differential expression.

📄 PDF Abstract BibTeX arXiv:1704.07008

Code (1)

wilsoncai1992/adaptest

Tasks

AstronomyMarketingTwo-sample testingvalidVocal Bursts Intensity Prediction

Similar Papers 제목 키워드 기반

Forecasting high-frequency financial time series: an adaptive learning approach with the order book data

2021-02-27 · Parley Ruogu Yang

This paper proposes a forecast-centric adaptive learning model that engages with the past studies on the order book and high-frequency data, with applications to hypothesis testing. In line with the past literature, we p…

Time SeriesTime Series Analysis

AdaPT-GMM: Powerful and robust covariate-assisted multiple testing

2021-06-30 · Patrick Chao, William Fithian

We propose a new empirical Bayes method for covariate-assisted multiple testing with false discovery rate (FDR) control, where we model the local false discovery rate for each hypothesis as a function of both its covaria…

A projection pursuit framework for testing general high-dimensional hypothesis

2017-05-02 · Yinchu Zhu, Jelena Bradic

This article develops a framework for testing general hypothesis in high-dimensional models where the number of variables may far exceed the number of observations. Existing literature has considered less than a handful …

Variable SelectionVocal Bursts Intensity Prediction

Large Deviation Analysis of Score-based Hypothesis Testing

2024-01-27 · Enmao Diao, Taposh Banerjee, Vahid Tarokh

Score-based statistical models play an important role in modern machine learning, statistics, and signal processing. For hypothesis testing, a score-based hypothesis test is proposed in \cite{wu2022score}. We analyze the…

Two-Sample Tests for Large Random Graphs Using Network Statistics

2017-05-17 · Debarghya Ghoshdastidar, Maurilio Gutzeit, Alexandra Carpentier, Ulrike Von Luxburg

We consider a two-sample hypothesis testing problem, where the distributions are defined on the space of undirected graphs, and one has access to only one observation from each model. A motivating example for this proble…

Two-sample testingVocal Bursts Valence Prediction