paper-with-me

홈 › Papers

Approximate Gaussianity Beyond Initialisation in Neural Networks

2025-10-06 · Edward Hirst, Sanjaye Ramgoolam arxiv

Ensembles of neural network weight matrices are studied through the training process for the MNIST classification problem, testing the efficacy of matrix models for representing their distributions, under assumptions of Gaussianity and permutation-symmetry. The general 13-parameter permutation invariant Gaussian matrix models are found to be effective models for the correlated Gaussianity in the weight matrices, beyond the range of applicability of the simple Gaussian with independent identically distributed matrix variables, and notably well beyond the initialisation step. The representation theoretic model parameters, and the graph-theoretic characterisation of the permutation invariant matrix observables give an interpretable framework for the best-fit model and for small departures from Gaussianity. Additionally, the Wasserstein distance is calculated for this class of models and used to quantify the movement of the distributions over training. Throughout the work, the effects of varied initialisation regimes, regularisation, layer depth, and layer width are tested for this formalism, identifying limits where particular departures from Gaussianity are enhanced and how more general, yet still highly-interpretable, models can be developed.

📄 PDF Abstract BibTeX arXiv:2510.05218

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Vector-valued self-normalized concentration inequalities beyond sub-Gaussianity

2025-11-05 · Diego Martinez-Taboada, Tomas Gonzalez, Aaditya Ramdas arxiv

The study of self-normalized processes plays a crucial role in a wide range of applications, from sequential decision-making to econometrics. While the behavior of self-normalized concentration has been widely investigat…

How Controlling the Variance can Improve Training Stability of Sparsely Activated DNNs and CNNs

2026-02-05 · Emily Dent, Jared Tanner arxiv

The Edge-of-Chaos (EoC) theory developed for the random initialization of deep networks allows more efficient training by both preserving information in the initial outputs of the network and minimising exploding or vani…

Gaussian Processes

Gaussianity and typicality in matrix distributional semantics

2019-12-19 · Sanjaye Ramgoolam, Mehrnoosh Sadrzadeh, Lewis Sword

Constructions in type-driven compositional distributional semantics associate large collections of matrices of size $D$ to linguistic corpora. We develop the proposal of analysing the statistical characteristics of this …

Graph neural network initialisation of quantum approximate optimisation

2021-11-04 · Nishant Jain, Brian Coyle, Elham Kashefi, Niraj Kumar

Approximate combinatorial optimisation has emerged as one of the most promising application areas for quantum computers, particularly those in the near term. In this work, we focus on the quantum approximate optimisation…

Graph Neural NetworkMeta-Learning

Testing distributional assumptions of learning algorithms

2022-04-14 · Ronitt Rubinfeld, Arsen Vasilyan

There are many high dimensional function classes that have fast agnostic learning algorithms when assumptions on the distribution of examples can be made, such as Gaussianity or uniformity over the domain. But how can on…