paper-with-me

홈 › Papers

Canonical Regularisation of Wide Feature-Learning Neural Networks

2026-05-18 · George Whittle, Pranav Vaidhyanathan, Juliusz Ziomek, Natalia Ares, Maike A. Osborne arxiv

Wide neural networks in the feature-learning regime drive modern deep learning, and yet they remain far less studied than their kernel-regime counterparts. We consider a critical yet under-explored difference between these two regimes: the regulariser and prior implied by gradient flow training. This canonical regularisation property is well-studied in kernel regime networks -- of all the infinite global minima, gradient flow selects exactly the vanishing ridge solution -- and underpins the celebrated NN-GP correspondence, precisely allowing the modelling of noise during training. However, we prove ridge regularisation biases gradient flow in feature-learning regime networks, even in the infinitesimal limit of vanishing regularisation. Over training, ridge distorts the inductive bias of the network, with a particular damage done to pretrained networks where the implicit prior is informative. We resolve this by axiomatising the canonical regulariser as a regime-agnostic function-space energy and lift, which uniquely identifies ridge in the kernel regime, and crucially generalises to the feature-learning regime. By studying the Riemannian geometry of feature-learning networks, we derive geodesic ridge from our framework, generalising ridge to the feature-learning regime. Correspondingly, we prove the canonical function-space prior is a Riemannian Gibbs Process, generalising the more familiar Gaussian Process. As a practical contribution, we propose arc ridge as a minimax-robust, scalable surrogate to geodesic ridge, revealing a deep relationship between early stopping and canonical regularisation across learning regimes. Finally, we demonstrate the consequences of our theory empirically on both image processing and NLP transfer-learning problems.

📄 PDF Abstract BibTeX arXiv:2605.18180

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Temporal Smoothness Regularisers for Neural Link Predictors

2023-09-16 · Manuel Dileo, Pasquale Minervini, Matteo Zignani, Sabrina Gaito

Most algorithms for representation learning and link prediction on relational data are designed for static data. However, the data to which they are applied typically evolves over time, including online social networks o…

Knowledge GraphsLink PredictionPredictionRecommendation Systems+1

Regularisation in neural networks: a survey and empirical analysis of approaches

2026-01-30 · Christiaan P. Opperman, Anna S. Bosman, Katherine M. Malan arxiv

Despite huge successes on a wide range of tasks, neural networks are known to sometimes struggle to generalise to unseen data. Many approaches have been proposed over the years to promote the generalisation ability of ne…

On the Interpretability of Regularisation for Neural Networks Through Model Gradient Similarity

2022-05-25 · Vincent Szolnoky, Viktor Andersson, Balazs Kulcsar, Rebecka Jörnsten

Most complex machine learning and modelling techniques are prone to over-fitting and may subsequently generalise poorly to future data. Artificial neural networks are no different in this regard and, despite having a lev…

On Sparsity in Overparametrised Shallow ReLU Networks

2020-06-18 · Jaume de Dios, Joan Bruna

The analysis of neural network training beyond their linearization regime remains an outstanding open question, even in the simplest setup of a single hidden-layer. The limit of infinitely wide networks provides an appea…

Open-Ended Question Answering

Implicit Regularisation in Diffusion Models: An Algorithm-Dependent Generalisation Analysis

2025-07-04 · Tyler Farghly, Patrick Rebeschini, George Deligiannidis, Arnaud Doucet arxiv

The success of denoising diffusion models raises important questions regarding their generalisation behaviour, particularly in high-dimensional settings. Notably, it has been shown that when training and sampling are per…