Provable More Data Hurt in High Dimensional Least Squares Estimator
This paper investigates the finite-sample prediction risk of the high-dimensional least squares estimator. We derive the central limit theorem for the prediction risk when both the sample size and the number of features tend to infinity. Furthermore, the finite-sample distribution and the confidence interval of the prediction risk are provided. Our theoretical results demonstrate the sample-wise nonmonotonicity of the prediction risk and confirm "more data hurt" phenomenon.
Code (0)
등록된 구현이 없습니다.
Tasks
PredictionVocal Bursts Intensity PredictionSimilar Papers 제목 키워드 기반
Cleansing & expanding the HURTLEX(el) with a multidimensional categorization of offensive words
We present a cleansed version of the multilingual lexicon HURTLEX-(EL) comprising 737 offensive words of Modern Greek. We worked bottom-up in two annotation rounds and developed detailed guidelines by cross-classifying w…
Provable Privacy Attacks on Trained Shallow Neural Networks
We study what provable privacy attacks can be shown on trained, 2-layer ReLU neural networks. We explore two types of attacks; data reconstruction attacks, and membership inference attacks. We prove that theoretical resu…
The Curious Case of Adversarially Robust Models: More Data Can Help, Double Descend, or Hurt Generalization
Adversarial training has shown its ability in producing models that are robust to perturbations on the input data, but usually at the expense of decrease in the standard accuracy. To mitigate this issue, it is commonly b…
ClassificationGeneral ClassificationPractical Data-Dependent Metric Compression with Provable Guarantees
We introduce a new distance-preserving compact representation of multi-dimensional point-sets. Given n points in a d-dimensional space where each coordinate is represented using B bits (i.e., dB bits per point), it produ…
QuantizationTime SeriesTime Series AnalysisWhy adversarial training can hurt robust accuracy
Machine learning classifiers with high test accuracy often perform poorly under adversarial attacks. It is commonly believed that adversarial training alleviates this issue. In this paper, we demonstrate that, surprising…