paper-with-me

홈 › Papers

Geometry and Optimization of Shallow Polynomial Networks

2025-01-10 · Yossi Arjevani, Joan Bruna, Joe Kileel, Elzbieta Polak, Matthew Trager

We study shallow neural networks with polynomial activations. The function space for these models can be identified with a set of symmetric tensors with bounded rank. We describe general features of these networks, focusing on the relationship between width and optimization. We then consider teacher-student problems, that can be viewed as a problem of low-rank tensor approximation with respect to a non-standard inner product that is induced by the data distribution. In this setting, we introduce a teacher-metric discriminant which encodes the qualitative behavior of the optimization as a function of the training data distribution. Finally, we focus on networks with quadratic activations, presenting an in-depth analysis of the optimization landscape. In particular, we present a variation of the Eckart-Young Theorem characterizing all critical points and their Hessian signatures for teacher-student problems with quadratic networks and Gaussian training data.

📄 PDF Abstract BibTeX arXiv:2501.06074

Code (0)

등록된 구현이 없습니다.

Methods 이 논문이 사용한 방법론

SET Dynamic Sparse Training method where weight mask is updated randomly periodically
Focus 설명 없음

Similar Papers 제목 키워드 기반

On the existence of global minima and convergence analyses for gradient descent methods in the training of deep neural networks

2021-12-17 · Arnulf Jentzen, Adrian Riekert

In this article we study fully-connected feedforward deep ReLU ANNs with an arbitrarily large number of hidden layers and we prove convergence of the risk of the GD optimization method with random initializations in the …

Shallow neural network representation of polynomials

2022-08-17 · Aleksandr Beknazaryan

We show that $d$-variate polynomials of degree $R$ can be represented on $[0,1]^d$ as shallow neural networks of width $2(R+d)^d$. Also, by SNN representation of localized Taylor polynomials of univariate $C^\beta$-smoot…

regression

Learning shallow quantum circuits

2024-01-18 · Hsin-Yuan Huang, Yunchao Liu, Michael Broughton, Isaac Kim 외

Despite fundamental interests in learning quantum circuits, the existence of a computationally efficient algorithm for learning shallow quantum circuits remains an open question. Because shallow quantum circuits can gene…

The Geometry of Polynomial Group Convolutional Neural Networks

2026-03-31 · Yacoub Hendi, Daniel Persson, Magdalena Larfors arxiv

We study polynomial group convolutional neural networks (PGCNNs) for an arbitrary finite group $G$. In particular, we introduce a new mathematical framework for PGCNNs using the language of graded group algebras. This fr…

SUPN: Shallow Universal Polynomial Networks

2025-11-26 · Zachary Morrow, Michael Penwarden, Brian Chen, Aurya Javeed 외 arxiv

Deep neural networks (DNNs) and Kolmogorov-Arnold networks (KANs) are popular methods for function approximation due to their flexibility and expressivity. However, they typically require a large number of trainable para…