paper-with-me

홈 › Papers

General Loss Functions Lead to (Approximate) Interpolation in High Dimensions

2023-03-13 · Kuo-Wei Lai, Vidya Muthukumar

We provide a unified framework that applies to a general family of convex losses across binary and multiclass settings in the overparameterized regime to approximately characterize the implicit bias of gradient descent in closed form. Specifically, we show that the implicit bias is approximated (but not exactly equal to) the minimum-norm interpolation in high dimensions, which arises from training on the squared loss. In contrast to prior work, which was tailored to exponentially-tailed losses and used the intermediate support-vector-machine formulation, our framework directly builds on the primal-dual analysis of Ji and Telgarsky (2021), allowing us to provide new approximate equivalences for general convex losses through a novel sensitivity analysis. Our framework also recovers existing exact equivalence results for exponentially-tailed losses across binary and multiclass settings. Finally, we provide evidence for the tightness of our techniques and use our results to demonstrate the effect of certain loss functions designed for out-of-distribution problems on the closed-form solution.

📄 PDF Abstract BibTeX arXiv:2303.07475

Code (0)

등록된 구현이 없습니다.

Tasks

Vocal Bursts Intensity Prediction

Similar Papers 제목 키워드 기반

On the Oracle Complexity of Interpolation-Based Gradient Descent

2026-06-18 · Dongmin Lee, William Lu, Anuran Makur arxiv

Recent work on first-order optimizers for empirical risk minimization (ERM) has suggested that smoothness of ERM loss functions in the training data, rather than in the optimization parameters, can be leveraged to improv…

Simplicity bias and optimization threshold in two-layer ReLU networks

2024-10-03 · Etienne Boursier, Nicolas Flammarion

Understanding generalization of overparametrized neural networks remains a fundamental challenge in machine learning. Most of the literature mostly studies generalization from an interpolation point of view, taking conve…

In-Context Learning

The limitation of neural nets for approximation and optimization

2023-11-21 · Tommaso Giovannelli, Oumaima Sohab, Luis Nunes Vicente

We are interested in assessing the use of neural networks as surrogate models to approximate and minimize objective functions in optimization problems. While neural networks are widely used for machine learning tasks suc…

regression

Gradient Descent Quantizes ReLU Network Features

2018-03-22 · Hartmut Maennel, Olivier Bousquet, Sylvain Gelly

Deep neural networks are often trained in the over-parametrized regime (i.e. with far more parameters than training examples), and understanding why the training converges to solutions that generalize remains an open pro…

Quantization

Fast Signal Interpolation Through Zero-padding and FFT/IFFT

2024-07-09 · Zijun Gong

Based on the sampling theorem, interpolation should be conducted by employing the sinc functions as the kernels. Inspired by the fact that the discrete Fourier transform (DFT) is sampled from the discrete time Fourier tr…