paper-with-me

홈 › Papers

Making Progress Based on False Discoveries

2022-04-19 · Roi Livni

The study of adaptive data analysis examines how many statistical queries can be answered accurately using a fixed dataset while avoiding false discoveries (statistically inaccurate answers). In this paper, we tackle a question that precedes the field of study: Is data only valuable when it provides accurate answers to statistical queries? To answer this question, we use Stochastic Convex Optimization as a case study. In this model, algorithms are considered as analysts who query an estimate of the gradient of a noisy function at each iteration and move towards its minimizer. It is known that $O(1/\epsilon^2)$ examples can be used to minimize the objective function, but none of the existing methods depend on the accuracy of the estimated gradients along the trajectory. Therefore, we ask: How many samples are needed to minimize a noisy convex function if we require $\epsilon$-accurate estimates of $O(1/\epsilon^2)$ gradients? Or, might it be that inaccurate gradient estimates are \emph{necessary} for finding the minimum of a stochastic convex function at an optimal statistical rate? We provide two partial answers to this question. First, we show that a general analyst (queries that may be maliciously chosen) requires $\Omega(1/\epsilon^3)$ samples, ruling out the possibility of a foolproof mechanism. Second, we show that, under certain assumptions on the oracle, $\tilde \Omega(1/\epsilon^{2.5})$ samples are necessary for gradient descent to interact with the oracle. Our results are in contrast to classical bounds that show that $O(1/\epsilon^2)$ samples can optimize the population risk to an accuracy of $O(\epsilon)$, but with spurious gradients.

📄 PDF Abstract BibTeX arXiv:2204.08809

Code (0)

등록된 구현이 없습니다.

Methods 이 논문이 사용한 방법론

NON 설명 없음

Similar Papers 제목 키워드 기반

False Discovery Proportion control for aggregated Knockoffs

2023-09-21 · NeurIPS 2023 11

Controlled variable selection is an important analytical step in various scientific fields, such as brain imaging or genomics. In these high-dimensional data settings, considering too many variables leads to poor models …

Communication-Efficient False Discovery Rate Control via Knockoff Aggregation

2015-06-17 · Weijie Su, Junyang Qian, Linxi Liu

The false discovery rate (FDR)---the expected fraction of spurious discoveries among all the discoveries---provides a popular statistical assessment of the reproducibility of scientific studies in various disciplines. In…

A New Perspective on Pool-Based Active Classification and False-Discovery Control

2020-08-14 · NeurIPS 2019 12 · Lalit Jain, Kevin Jamieson

In many scientific settings there is a need for adaptive experimental design to guide the process of identifying regions of the search space that contain as many true positives as possible subject to a low rate of false …

Active LearningBinary ClassificationClassificationExperimental Design+1

TriSig: Assessing the statistical significance of triclusters

2023-06-01 · Leonardo Alexandre, Rafael S. Costa, Rui Henriques

Tensor data analysis allows researchers to uncover novel patterns and relationships that cannot be obtained from matrix data alone. The information inferred from the patterns provides valuable insights into disease progr…

False Discovery Rate Control and Statistical Quality Assessment of Annotators in Crowdsourced Ranking

2016-05-19 · Qianqian Xu, Jiechao Xiong, Xiaochun Cao, Yuan YAO

With the rapid growth of crowdsourcing platforms it has become easy and relatively inexpensive to collect a dataset labeled by multiple annotators in a short time. However due to the lack of control over the quality of t…

PositionSociology