paper-with-me

홈 › Papers

Provable More Data Hurt in High Dimensional Least Squares Estimator

2020-08-14 · Zeng Li, Chuanlong Xie, Qinwen Wang

This paper investigates the finite-sample prediction risk of the high-dimensional least squares estimator. We derive the central limit theorem for the prediction risk when both the sample size and the number of features tend to infinity. Furthermore, the finite-sample distribution and the confidence interval of the prediction risk are provided. Our theoretical results demonstrate the sample-wise nonmonotonicity of the prediction risk and confirm "more data hurt" phenomenon.

📄 PDF Abstract BibTeX arXiv:2008.06296

Code (0)

등록된 구현이 없습니다.

Tasks

PredictionVocal Bursts Intensity Prediction

Similar Papers 제목 키워드 기반

Cleansing & expanding the HURTLEX(el) with a multidimensional categorization of offensive words

2022-07-01 · NAACL (WOAH) 2022 7 · Vivian Stamou, Iakovi Alexiou, Antigone Klimi, Eleftheria Molou 외

We present a cleansed version of the multilingual lexicon HURTLEX-(EL) comprising 737 offensive words of Modern Greek. We worked bottom-up in two annotation rounds and developed detailed guidelines by cross-classifying w…

Provable Privacy Attacks on Trained Shallow Neural Networks

2024-10-10 · Guy Smorodinsky, Gal Vardi, Itay Safran

We study what provable privacy attacks can be shown on trained, 2-layer ReLU neural networks. We explore two types of attacks; data reconstruction attacks, and membership inference attacks. We prove that theoretical resu…

The Curious Case of Adversarially Robust Models: More Data Can Help, Double Descend, or Hurt Generalization

2020-02-25 · Yifei Min, Lin Chen, Amin Karbasi

Adversarial training has shown its ability in producing models that are robust to perturbations on the input data, but usually at the expense of decrease in the standard accuracy. To mitigate this issue, it is commonly b…

ClassificationGeneral Classification

Practical Data-Dependent Metric Compression with Provable Guarantees

2017-12-01 · NeurIPS 2017 12 · Piotr Indyk, Ilya Razenshteyn, Tal Wagner

We introduce a new distance-preserving compact representation of multi-dimensional point-sets. Given n points in a d-dimensional space where each coordinate is represented using B bits (i.e., dB bits per point), it produ…

QuantizationTime SeriesTime Series Analysis

Why adversarial training can hurt robust accuracy

2022-03-03 · Jacob Clarysse, Julia Hörmann, Fanny Yang

Machine learning classifiers with high test accuracy often perform poorly under adversarial attacks. It is commonly believed that adversarial training alleviates this issue. In this paper, we demonstrate that, surprising…