paper-with-me

홈 › Papers

Module-wise Training of Neural Networks via the Minimizing Movement Scheme

2023-09-29 · NeurIPS 2023 11

Greedy layer-wise or module-wise training of neural networks is compelling in constrained and on-device settings where memory is limited, as it circumvents a number of problems of end-to-end back-propagation. However, it suffers from a stagnation problem, whereby early layers overfit and deeper layers stop increasing the test accuracy after a certain depth. We propose to solve this issue by introducing a module-wise regularization inspired by the minimizing movement scheme for gradient flows in distribution space. We call the method TRGL for Transport Regularized Greedy Learning and study it theoretically, proving that it leads to greedy modules that are regular and that progressively solve the task. Experimentally, we show improved accuracy of module-wise training of various architectures such as ResNets, Transformers and VGG, when our regularization is added, superior to that of other module-wise training methods and often to end-to-end training, with as much as 60% less memory usage.

📄 PDF Abstract BibTeX arXiv:2309.17357

Code (0)

등록된 구현이 없습니다.

Methods 이 논문이 사용한 방법론

Convolution A convolution is a type of matrix operation, consisting of a kernel, a small matrix of weights, that slides over input data performing element-wise multiplication with the…
Dense Connections Dense Connections, or Fully Connected Connections, are a type of layer in a deep neural network that use a linear operation where every input is connected to every output…
Dropout Dropout is a regularization technique for neural networks that drops a unit (along with connections) at training time with a specified probability $p$ (a common value is…
Softmax The Softmax output function transforms a previous layer's output into a vector of probabilities. It is commonly used for multiclass classification. Given an input vector $x$…
Max Pooling Max Pooling is a pooling operation that calculates the maximum value for patches of a feature map, and uses it to create a downsampled (pooled) feature map. It is usually…

Similar Papers 제목 키워드 기반

Block-wise Training of Residual Networks via the Minimizing Movement Scheme

2022-10-03 · Skander Karkar, Ibrahim Ayed, Emmanuel de Bézenac, Patrick Gallinari

End-to-end backpropagation has a few shortcomings: it requires loading the entire model during training, which can be impossible in constrained settings, and suffers from three locking problems (forward locking, update l…

A mean curvature flow arising in adversarial training

2024-04-22 · Leon Bungert, Tim Laux, Kerrek Stinson

We connect adversarial training for binary classification to a geometric evolution equation for the decision boundary. Relying on a perspective that recasts adversarial training as a regularization problem, we introduce …

Binary Classification

Gradient Flow Based Phase-Field Modeling Using Separable Neural Networks

2024-05-09 · Revanth Mattey, Susanta Ghosh

The $L^2$ gradient flow of the Ginzburg-Landau free energy functional leads to the Allen Cahn equation that is widely used for modeling phase separation. Machine learning methods for solving the Allen-Cahn equation in it…

Tensor Decomposition

Semi-discrete optimization through semi-discrete optimal transport: a framework for neural architecture search

2020-06-26 · Nicolas Garcia Trillos, Javier Morales

In this paper we introduce a theoretical framework for semi-discrete optimization using ideas from optimal transport. Our primary motivation is in the field of deep learning, and specifically in the task of neural archit…

Neural Architecture Search

Decoding Human Activities: Analyzing Wearable Accelerometer and Gyroscope Data for Activity Recognition

2023-10-03 · Utsab Saha, Sawradip Saha, Tahmid Kabir, Shaikh Anowarul Fattah 외

A person's movement or relative positioning can be effectively captured by different types of sensors and corresponding sensor output can be utilized in various manipulative techniques for the classification of different…

Activity RecognitionHuman Activity Recognition