Implicit regularization for deep neural networks driven by an Ornstein-Uhlenbeck like process
We consider networks, trained via stochastic gradient descent to minimize $\ell_2$ loss, with the training labels perturbed by independent noise at each iteration. We characterize the behavior of the training dynamics near any parameter vector that achieves zero training error, in terms of an implicit regularization term corresponding to the sum over the data points, of the squared $\ell_2$ norm of the gradient of the model with respect to the parameter vector, evaluated at each data point. This holds for networks of any connectivity, width, depth, and choice of activation function. We interpret this implicit regularization term for three simple settings: matrix sensing, two layer ReLU networks trained on one-dimensional data, and two layer networks with sigmoid activations trained on a single datapoint. For these settings, we show why this new and general implicit regularization effect drives the networks towards "simple" models.
Code (0)
등록된 구현이 없습니다.
Methods 이 논문이 사용한 방법론
Similar Papers 제목 키워드 기반
Exact simulation scheme for the Ornstein-Uhlenbeck driven stochastic volatility model with the Karhunen-Loève expansions
This study proposes a new exact simulation scheme of the Ornstein-Uhlenbeck driven stochastic volatility model. With the Karhunen-Lo\`eve expansions, the stochastic volatility path following the Ornstein-Uhlenbeck proces…
Pricing Energy Derivatives in Markets Driven by Tempered Stable and CGMY Processes of Ornstein-Uhlenbeck Type
In this study we consider the pricing of energy derivatives when the evolution of spot prices follows a tempered stable or a CGMY driven Ornstein- Uhlenbeck process. To this end, we first calculate the characteristic fun…
Estimation of Ornstein-Uhlenbeck Process Using Ultra-High-Frequency Data with Application to Intraday Pairs Trading Strategy
When stock prices are observed at high frequencies, more information can be utilized in estimation of parameters of the price process. However, high-frequency data are contaminated by the market microstructure noise whic…
parameter estimationSparse inference of the drift of a high-dimensional Ornstein-Uhlenbeck process
Given the observation of a high-dimensional Ornstein-Uhlenbeck (OU) process in continuous time, we proceed to the inference of the drift parameter under a row-sparsity assumption. Towards that aim, we consider the negati…
Variable SelectionValuing Exchange Options Under an Ornstein-Uhlenbeck Covariance Model
In this paper we study the pricing of exchange options under a dynamic described by stochastic correlation with random jumps. In particular, we consider a Ornstein-Uhlenbeck covariance model with Levy Background Noise Pr…