paper-with-me

Papers

On discretisation drift and smoothness regularisation in neural network training

2023-10-21 · Mihaela Claudia Rosca

The deep learning recipe of casting real-world problems as mathematical optimisation and tackling the optimisation by training deep neural networks using gradient-based optimisation has undoubtedly proven to be a fruitful one. The understanding behind why deep learning works, however, has lagged behind its practical significance. We aim to make steps towards an improved understanding of deep learning with a focus on optimisation and model regularisation. We start by investigating gradient descent (GD), a discrete-time algorithm at the basis of most popular deep learning optimisation algorithms. Understanding the dynamics of GD has been hindered by the presence of discretisation drift, the numerical integration error between GD and its often studied continuous-time counterpart, the negative gradient flow (NGF). To add to the toolkit available to study GD, we derive novel continuous-time flows that account for discretisation drift. Unlike the NGF, these new flows can be used to describe learning rate specific behaviours of GD, such as training instabilities observed in supervised learning and two-player games. We then translate insights from continuous time into mitigation strategies for unstable GD dynamics, by constructing novel learning rate schedules and regularisers that do not require additional hyperparameters. Like optimisation, smoothness regularisation is another pillar of deep learning's success with wide use in supervised learning and generative modelling. Despite their individual significance, the interactions between smoothness regularisation and optimisation have yet to be explored. We find that smoothness regularisation affects optimisation across multiple deep learning domains, and that incorporating smoothness regularisation in reinforcement learning leads to a performance boost that can be recovered using adaptions to optimisation methods.

📄 PDF Abstract BibTeX arXiv:2310.14036

Code (0)

등록된 구현이 없습니다.

Tasks

Deep LearningNumerical Integration

Methods 이 논문이 사용한 방법론

Focus 설명 없음

Similar Papers 제목 키워드 기반

Implicit Regularisation in Diffusion Models: An Algorithm-Dependent Generalisation Analysis

2025-07-04 · Tyler Farghly, Patrick Rebeschini, George Deligiannidis, Arnaud Doucet arxiv

The success of denoising diffusion models raises important questions regarding their generalisation behaviour, particularly in high-dimensional settings. Notably, it has been shown that when training and sampling are per…

Implicit regularisation in stochastic gradient descent: from single-objective to two-player games

2023-07-11 · Mihaela Rosca, Marc Peter Deisenroth

Recent years have seen many insights on deep learning optimisation being brought forward by finding implicit regularisation effects of commonly used gradient-based optimisers. Understanding implicit regularisation can no…

TASER: Task-Aware Stein Regularisation for Geometry-Driven Robustness

2026-05-28 · Michał Kozyra, Gesine Reinert arxiv

Modern deep networks remain fragile under distribution shift and adversarial perturbations, often due to excessive or poorly structured input sensitivity. We introduce TASER (Task-Aware Stein Regularisation), a training-…

Adversarial Robustness

Using Deep Image Prior to Assist Variational Selective Segmentation Deep Learning Algorithms

2021-12-01 · Liam Burrows, Ke Chen, Francesco Torella

Variational segmentation algorithms require a prior imposed in the form of a regularisation term to enforce smoothness of the solution. Recently, it was shown in the Deep Image Prior work that the explicit regularisation…

Non-asymptotic bounds for sampling algorithms without log-concavity

2018-08-21 · Mateusz B. Majka, Aleksandar Mijatović, Lukasz Szpruch

Discrete time analogues of ergodic stochastic differential equations (SDEs) are one of the most popular and flexible tools for sampling high-dimensional probability measures. Non-asymptotic analysis in the $L^2$ Wasserst…