paper-with-me

Papers

Unification of Symmetries Inside Neural Networks: Transformer, Feedforward and Neural ODE

2024-02-04 · Koji Hashimoto, Yuji Hirono, Akiyoshi Sannai

Understanding the inner workings of neural networks, including transformers, remains one of the most challenging puzzles in machine learning. This study introduces a novel approach by applying the principles of gauge symmetries, a key concept in physics, to neural network architectures. By regarding model functions as physical observables, we find that parametric redundancies of various machine learning models can be interpreted as gauge symmetries. We mathematically formulate the parametric redundancies in neural ODEs, and find that their gauge symmetries are given by spacetime diffeomorphisms, which play a fundamental role in Einstein's theory of gravity. Viewing neural ODEs as a continuum version of feedforward neural networks, we show that the parametric redundancies in feedforward neural networks are indeed lifted to diffeomorphisms in neural ODEs. We further extend our analysis to transformer models, finding natural correspondences with neural ODEs and their gauge symmetries. The concept of gauge symmetries sheds light on the complex behavior of deep learning models through physics and provides us with a unifying perspective for analyzing various machine learning architectures.

📄 PDF Abstract BibTeX arXiv:2402.02362

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Every Feedforward Neural Network Definable in an o-Minimal Structure Has Finite Sample Complexity

2026-05-08 · Anastasis Kratsios, Gregory Cousins, Haitz Sáez de Ocáriz Borde, Bum Jun Kim 외 arxiv

We show that, in a precise sense, a broad class of feedforward neural networks learn (have finite sample complexity) in the PAC model: every fixed finite feedforward architecture whose layers are definable in an o-minima…

Model Compression via Symmetries of the Parameter Space

2021-09-29 · Iordan Ganev, Robin Walters

We provide a theoretical framework for neural networks in terms of the representation theory of quivers, thus revealing symmetries of the parameter space of neural networks. An exploitation of these symmetries leads to a…

modelModel Compression

Hidden symmetries of ReLU networks

2023-06-09 · J. Elisenda Grigsby, Kathryn Lindsey, David Rolnick

The parameter space for any fixed architecture of feedforward ReLU neural networks serves as a proxy during training for the associated class of functions - but how faithful is this representation? It is known that many …

Global Structure-from-Motion Meets Feedforward Reconstruction

2026-05-25 · Linfei Pan, Johannes Schönberger, Marc Pollefeys arxiv

Structure-from-Motion -- the process of simultaneously estimating camera poses and 3D scene structure from a collection of images -- remains a central challenge in computer vision, with many open problems yet to be solve…

3D Reconstruction

Should attention be all we need? The epistemic and ethical implications of unification in machine learning

2022-05-09 · Nic Fishman, Leif Hancox-Li

"Attention is all you need" has become a fundamental precept in machine learning research. Originally designed for machine translation, transformers and the attention mechanisms that underpin them now find success across…

AllBIG-bench Machine LearningDiversityMachine Translation