paper-with-me

Papers

LION-DG: Layer-Informed Initialization with Deep Gradient Protocols for Accelerated Neural Network Training

2026-01-05 · Hyunjun Kim arxiv

Weight initialization remains decisive for neural network optimization, yet existing methods are largely layer-agnostic. We study initialization for deeply-supervised architectures with auxiliary classifiers, where untrained auxiliary heads can destabilize early training through gradient interference. We propose LION-DG, a layer-informed initialization that zero-initializes auxiliary classifier heads while applying standard He-initialization to the backbone. We prove that this implements Gradient Awakening: auxiliary gradients are exactly zero at initialization, then phase in naturally as weights grow -- providing an implicit warmup without hyperparameters. Experiments on CIFAR-10 and CIFAR-100 with DenseNet-DS and ResNet-DS architectures demonstrate: (1) DenseNet-DS: +8.3% faster convergence on CIFAR-10 with comparable accuracy, (2) Hybrid approach: Combining LSUV with LION-DG achieves best accuracy (81.92% on CIFAR-10), (3) ResNet-DS: Positive speedup on CIFAR-100 (+11.3%) with side-tap auxiliary design. We identify architecture-specific trade-offs and provide clear guidelines for practitioners. LION-DG is simple, requires zero hyperparameters, and adds no computational overhead.

📄 PDF Abstract BibTeX arXiv:2601.02105

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Are Two Hidden Layers Still Enough for the Physics-Informed Neural Networks?

2024-12-26 · Vasiliy A. Es'kin, Alexey O. Malkhanov, Mikhail E. Smorkalov

The article discusses the development of various methods and techniques for initializing and training neural networks with a single hidden layer, as well as training a separable physics-informed neural network consisting…

Initialization of Large Language Models via Reparameterization to Mitigate Loss Spikes

2024-10-07 · Kosuke Nishida, Kyosuke Nishida, Kuniko Saito

Loss spikes, a phenomenon in which the loss value diverges suddenly, is a fundamental issue in the pre-training of large language models. This paper supposes that the non-uniformity of the norm of the parameters is one o…

Data-driven Weight Initialization with Sylvester Solvers

2021-05-02 · Debasmit Das, Yash Bhalgat, Fatih Porikli

In this work, we propose a data-driven scheme to initialize the parameters of a deep neural network. This is in contrast to traditional approaches which randomly initialize parameters by sampling from transformed standar…

About rectified sigmoid function for enhancing the accuracy of Physics-Informed Neural Networks

2024-12-30 · Vasiliy A. Es'kin, Alexey O. Malkhanov, Mikhail E. Smorkalov

The article is devoted to the study of neural networks with one hidden layer and a modified activation function for solving physical problems. A rectified sigmoid activation function has been proposed to solve physical p…

Locally adaptive activation functions with slope recovery term for deep and physics-informed neural networks

2019-09-25 · Ameya D. Jagtap, Kenji Kawaguchi, George Em. Karniadakis

We propose two approaches of locally adaptive activation functions namely, layer-wise and neuron-wise locally adaptive activation functions, which improve the performance of deep and physics-informed neural networks. The…

Data Augmentation