paper-with-me

홈 › Papers

On the Lipschitz Constant of Deep Networks and Double Descent

2023-01-28 · Matteo Gamba, Hossein Azizpour, Mårten Björkman

Existing bounds on the generalization error of deep networks assume some form of smooth or bounded dependence on the input variable, falling short of investigating the mechanisms controlling such factors in practice. In this work, we present an extensive experimental study of the empirical Lipschitz constant of deep networks undergoing double descent, and highlight non-monotonic trends strongly correlating with the test error. Building a connection between parameter-space and input-space gradients for SGD around a critical point, we isolate two important factors -- namely loss landscape curvature and distance of parameters from initialization -- respectively controlling optimization dynamics around a critical point and bounding model function complexity, even beyond the training data. Our study presents novels insights on implicit regularization via overparameterization, and effective model complexity for networks trained in practice.

📄 PDF Abstract BibTeX arXiv:2301.12309

Code (1)

magamba/overparameterization 공식 구현 pytorch

Methods 이 논문이 사용한 방법론

Test 설명 없음
SGD Stochastic Gradient Descent is an iterative optimization technique that uses minibatches of data to form an expectation of the gradient, rather than the full gradient using…

Similar Papers 제목 키워드 기반

Concavifiability and convergence: necessary and sufficient conditions for gradient descent analysis

2019-05-28 · Thulasi Tholeti, Sheetal Kalyani

Convergence of the gradient descent algorithm has been attracting renewed interest due to its utility in deep learning applications. Even as multiple variants of gradient descent were proposed, the assumption that the gr…

Beyond Uniform Lipschitz Condition in Differentially Private Optimization

2022-06-21 · Rudrajit Das, Satyen Kale, Zheng Xu, Tong Zhang 외

Most prior results on differentially private stochastic gradient descent (DP-SGD) are derived under the simplistic assumption of uniform Lipschitzness, i.e., the per-sample gradients are uniformly bounded. We generalize …

Benchmarkingregression

Adaptive Sampling Distributed Stochastic Variance Reduced Gradient for Heterogeneous Distributed Datasets

2020-02-20 · Ilqar Ramazanli, Han Nguyen, Hai Pham, Sashank J. Reddi 외

We study distributed optimization algorithms for minimizing the average of \emph{heterogeneous} functions distributed across several machines with a focus on communication efficiency. In such settings, naively using the …

Distributed Optimization

On the Double Descent of Random Features Models Trained with SGD

2021-10-13 · Fanghui Liu, Johan A. K. Suykens, Volkan Cevher

We study generalization properties of random features (RF) regression in high dimensions optimized by stochastic gradient descent (SGD) in under-/over-parameterized regime. In this work, we derive precise non-asymptotic …

regression

Some Fundamental Aspects about Lipschitz Continuity of Neural Networks

2023-02-21 · Grigory Khromov, Sidak Pal Singh

Lipschitz continuity is a crucial functional property of any predictive model, that naturally governs its robustness, generalisation, as well as adversarial vulnerability. Contrary to other works that focus on obtaining …