Can we globally optimize cross-validation loss? Quasiconvexity in ridge regression
Models like LASSO and ridge regression are extensively used in practice due to their interpretability, ease of use, and strong theoretical guarantees. Cross-validation (CV) is widely used for hyperparameter tuning in these models, but do practical optimization methods minimize the true out-of-sample loss? A recent line of research promises to show that the optimum of the CV loss matches the optimum of the out-of-sample loss (possibly after simple corrections). It remains to show how tractable it is to minimize the CV loss. In the present paper, we show that, in the case of ridge regression, the CV loss may fail to be quasiconvex and thus may have multiple local optima. We can guarantee that the CV loss is quasiconvex in at least one case: when the spectrum of the covariate matrix is nearly flat and the noise in the observed responses is not too high. More generally, we show that quasiconvexity status is independent of many properties of the observed data (response norm, covariate-matrix right singular vectors and singular-value scaling) and has a complex dependence on the few that remain. We empirically confirm our theory using simulated experiments.
Code (0)
등록된 구현이 없습니다.
Tasks
regressionSimilar Papers 제목 키워드 기반
Decomposable sums and their implications on naturally quasiconvex risk measures
Convexity and quasiconvexity are two properties that capture the concept of diversification for risk measures. Between the two, there is natural quasiconvexity, an old but not so well-known property weaker than convexity…
Evaluation and Optimization of Leave-one-out Cross-validation for the Lasso
I develop an algorithm to produce the piecewise quadratic that computes leave-one-out cross-validation for the lasso as a function of its hyperparameter. The algorithm can be used to find exact hyperparameters that optim…
Bauer's Maximum Principle for Quasiconvex Functions
This note shows that in Bauer's maximum principle, the assumed convexity of the objective function can be relaxed to quasiconvexity.
Physics-Consistent Neural Networks for Learning Deformation and Director Fields in Microstructured Media with Loss-Based Validation Criteria
In this work, we study the mechanical behavior of solids with microstructure using the framework of Cosserat elasticity with a single unit director. This formulation captures the coupling between deformation and orientat…
Dynamical loss functions shape landscape topography and improve learning in artificial neural networks
Dynamical loss functions are derived from standard loss functions used in supervised classification tasks, but they are modified such that the contribution from each class periodically increases and decreases. These osci…