paper-with-me

홈 › Papers

Flat Minima in Linear Estimation and an Extended Gauss Markov Theorem

2023-11-18 · Simon Segert

We consider the problem of linear estimation, and establish an extension of the Gauss-Markov theorem, in which the bias operator is allowed to be non-zero but bounded with respect to a matrix norm of Schatten type. We derive simple and explicit formulas for the optimal estimator in the cases of Nuclear and Spectral norms (with the Frobenius case recovering ridge regression). Additionally, we analytically derive the generalization error in multiple random matrix ensembles, and compare with Ridge regression. Finally, we conduct an extensive simulation study, in which we show that the cross-validated Nuclear and Spectral regressors can outperform Ridge in several circumstances.

📄 PDF Abstract BibTeX arXiv:2311.11093

Code (1)

simonsegert/specreg 공식 구현

Tasks

regression

Similar Papers 제목 키워드 기반

Minimax Quantile Lower Bounds for Interactive Statistical Decision Making with Privacy

2026-06-22 · Raghav Bongole, Amirreza Zamani, Tobias J. Oechtering, Mikael Skoglund arxiv

Minimax risk and regret are expectation-based criteria and do not capture rare but consequential failures. To address this concern, we develop a $δ$-explicit minimax-quantile theory for interactive statistical decision m…

Decision Making

Flat minima generalize for low-rank matrix recovery

2022-03-07 · Lijun Ding, Dmitriy Drusvyatskiy, Maryam Fazel, Zaid Harchaoui

Empirical evidence suggests that for a variety of overparameterized nonlinear models, most notably in neural network training, the growth of the loss around a minimizer strongly impacts its performance. Flat minima -- th…

Matrix Completion

Wide flat minima and optimal generalization in classifying high-dimensional Gaussian mixtures

2020-10-27 · Carlo Baldassi, Enrico M. Malatesta, Matteo Negri, Riccardo Zecchina

We analyze the connection between minimizers with good generalizing properties and high local entropy regions of a threshold-linear classifier in Gaussian mixtures with the mean squared error loss function. We show that …

Unique Properties of Flat Minima in Deep Networks

2020-02-11 · Rotem Mulayoff, Tomer Michaeli

It is well known that (stochastic) gradient descent has an implicit bias towards flat minima. In deep neural network training, this mechanism serves to screen out minima. However, the precise effect that this has on the …

Dynamic Term Structure Models with Nonlinearities using Gaussian Processes

2023-05-18 · Tomasz Dubiel-Teleszynski, Konstantinos Kalogeropoulos, Nikolaos Karouzakis

The importance of unspanned macroeconomic variables for Dynamic Term Structure Models has been intensively discussed in the literature. To our best knowledge the earlier studies considered only linear interactions betwee…

Gaussian ProcessesPortfolio Optimization