paper-with-me

홈 › Papers

Asymptotic Characterisation of Robust Empirical Risk Minimisation Performance in the Presence of Outliers

2023-05-30 · Matteo Vilucchio, Emanuele Troiani, Vittorio Erba, Florent Krzakala

We study robust linear regression in high-dimension, when both the dimension $d$ and the number of data points $n$ diverge with a fixed ratio $\alpha=n/d$, and study a data model that includes outliers. We provide exact asymptotics for the performances of the empirical risk minimisation (ERM) using $\ell_2$-regularised $\ell_2$, $\ell_1$, and Huber losses, which are the standard approach to such problems. We focus on two metrics for the performance: the generalisation error to similar datasets with outliers, and the estimation error of the original, unpolluted function. Our results are compared with the information theoretic Bayes-optimal estimation bound. For the generalization error, we find that optimally-regularised ERM is asymptotically consistent in the large sample complexity limit if one perform a simple calibration, and compute the rates of convergence. For the estimation error however, we show that due to a norm calibration mismatch, the consistency of the estimator requires an oracle estimate of the optimal norm, or the presence of a cross-validation set not corrupted by the outliers. We examine in detail how performance depends on the loss function and on the degree of outlier corruption in the training set and identify a region of parameters where the optimal performance of the Huber loss is identical to that of the $\ell_2$ loss, offering insights into the use cases of different loss functions.

📄 PDF Abstract BibTeX arXiv:2305.18974

Code (1)

idephics/robustregression 공식 구현

Methods 이 논문이 사용한 방법론

Huber loss The Huber loss function describes the penalty incurred by an estimation procedure f. Huber (1964) defines the loss function piecewise by[1] L δ ( a ) = { 1 2 a 2 for | a |…
Linear Regression Linear Regression is a method for modelling a relationship between a dependent variable and independent variables. These models can be fit with numerous approaches. The most…
Focus 설명 없음

Similar Papers 제목 키워드 기반

Learning curves for the multi-class teacher-student perceptron

2022-03-22 · Elisabetta Cornacchia, Francesca Mignacco, Rodrigo Veiga, Cédric Gerbelot 외

One of the most classical results in high-dimensional learning theory provides a closed-form expression for the generalisation error of binary classification with the single-layer teacher-student perceptron on i.i.d. Gau…

Binary ClassificationLearning TheoryMulti-class Classification

Proper-Composite Loss Functions in Arbitrary Dimensions

2019-02-19 · Zac Cranko, Robert C. Williamson, Richard Nock

The study of a machine learning problem is in many ways is difficult to separate from the study of the loss function being used. One avenue of inquiry has been to look at these loss functions in terms of their properties…

Density Estimationscoring rule

Strong replica symmetry for high-dimensional disordered log-concave Gibbs measures

2020-09-27 · Jean Barbier, Dmitry Panchenko, Manuel Sáenz

We consider a generic class of log-concave, possibly random, (Gibbs) measures. We prove the concentration of an infinite family of order parameters called multioverlaps. Because they completely parametrise the quenched G…

Vocal Bursts Intensity Prediction

Optimal convex $M$-estimation via score matching

2024-03-25 · Oliver Y. Feng, Yu-Chun Kao, Min Xu, Richard J. Samworth

In the context of linear regression, we construct a data-driven convex loss function with respect to which empirical risk minimisation yields optimal asymptotic variance in the downstream estimation of the regression coe…

regression

Learning Gaussian Mixtures with Generalised Linear Models: Precise Asymptotics in High-dimensions

2021-06-07 · Bruno Loureiro, Gabriele Sicuro, Cédric Gerbelot, Alessandro Pacco 외

Generalised linear models for multi-class classification problems are one of the fundamental building blocks of modern machine learning tasks. In this manuscript, we characterise the learning of a mixture of $K$ Gaussian…

ClassificationMulti-class ClassificationVocal Bursts Intensity Prediction