paper-with-me

홈 › Papers

Consistent regression when oblivious outliers overwhelm

2020-09-30 · Tommaso d'Orsi, Gleb Novikov, David Steurer

We consider a robust linear regression model $y=X\beta^* + \eta$, where an adversary oblivious to the design $X\in \mathbb{R}^{n\times d}$ may choose $\eta$ to corrupt all but an $\alpha$ fraction of the observations $y$ in an arbitrary way. Prior to our work, even for Gaussian $X$, no estimator for $\beta^*$ was known to be consistent in this model except for quadratic sample size $n \gtrsim (d/\alpha)^2$ or for logarithmic inlier fraction $\alpha\ge 1/\log n$. We show that consistent estimation is possible with nearly linear sample size and inverse-polynomial inlier fraction. Concretely, we show that the Huber loss estimator is consistent for every sample size $n= \omega(d/\alpha^2)$ and achieves an error rate of $O(d/\alpha^2n)^{1/2}$. Both bounds are optimal (up to constant factors). Our results extend to designs far beyond the Gaussian case and only require the column span of $X$ to not contain approximately sparse vectors). (similar to the kind of assumption commonly made about the kernel space for compressed sensing). We provide two technically similar proofs. One proof is phrased in terms of strong convexity, extending work of [Tsakonas et al.'14], and particularly short. The other proof highlights a connection between the Huber loss estimator and high-dimensional median computations. In the special case of Gaussian designs, this connection leads us to a strikingly simple algorithm based on computing coordinate-wise medians that achieves optimal guarantees in nearly-linear time, and that can exploit sparsity of $\beta^*$. The model studied here also captures heavy-tailed noise distributions that may not even have a first moment.

📄 PDF Abstract BibTeX arXiv:2009.14774

Code (0)

등록된 구현이 없습니다.

Tasks

compressed sensingregression

Methods 이 논문이 사용한 방법론

Huber loss The Huber loss function describes the penalty incurred by an estimation procedure f. Huber (1964) defines the loss function piecewise by[1] L δ ( a ) = { 1 2 a 2 for | a |…
Linear Regression Linear Regression is a method for modelling a relationship between a dependent variable and independent variables. These models can be fit with numerous approaches. The most…

Similar Papers 제목 키워드 기반

Noise Statistics Oblivious GARD For Robust Regression With Sparse Outliers

2018-09-19 · Sreejith Kallummil, Sheetal Kalyani

Linear regression models contaminated by Gaussian noise (inlier) and possibly unbounded sparse outliers are common in many signal processing applications. Sparse recovery inspired robust regression (SRIRR) techniques are…

regression

Consistent Estimation for PCA and Sparse Regression with Oblivious Outliers

2021-11-04 · NeurIPS 2021 12 · Tommaso d'Orsi, Chih-Hung Liu, Rajai Nasser, Gleb Novikov 외

We develop machinery to design efficiently computable and consistent estimators, achieving estimation error approaching zero as the number of observations grows, when facing an oblivious adversary that may corrupt respon…

Matrix Completionregression

ROIDS: Robust Outlier-Aware Informed Down-Sampling

2026-01-27 · Alina Geiger, Martin Briesch, Dominik Sobania, Franz Rothlauf arxiv

Informed down-sampling (IDS) is known to improve performance in symbolic regression when combined with various selection strategies, especially tournament selection. However, recent work found that IDS's gains are not co…

Adaptive Hard Thresholding for Near-optimal Consistent Robust Regression

2019-03-19 · Arun Sai Suggala, Kush Bhatia, Pradeep Ravikumar, Prateek Jain

We study the problem of robust linear regression with response variable corruptions. We consider the oblivious adversary model, where the adversary corrupts a fraction of the responses in complete ignorance of the data. …

regression

Distribution-Independent Regression for Generalized Linear Models with Oblivious Corruptions

2023-09-20 · Ilias Diakonikolas, Sushrut Karmalkar, Jongho Park, Christos Tzamos

We demonstrate the first algorithms for the problem of regression for generalized linear models (GLMs) in the presence of additive oblivious noise. We assume we have sample access to examples $(x, y)$ where $y$ is a nois…

regression