paper-with-me

홈 › Papers

On Dissipativity of Cross-Entropy Loss in Training ResNets

2024-05-29 · Jens Püttschneider, Timm Faulwasser

The training of ResNets and neural ODEs can be formulated and analyzed from the perspective of optimal control. This paper proposes a dissipative formulation of the training of ResNets and neural ODEs for classification problems by including a variant of the cross-entropy as a regularization in the stage cost. Based on the dissipative formulation of the training, we prove that the trained ResNet exhibit the turnpike phenomenon. We then illustrate that the training exhibits the turnpike phenomenon by training on the two spirals and MNIST datasets. This can be used to find very shallow networks suitable for a given classification task.

📄 PDF Abstract BibTeX arXiv:2405.19013

Code (0)

등록된 구현이 없습니다.

Methods 이 논문이 사용한 방법론

Kaiming Initialization 설명 없음
Max Pooling Max Pooling is a pooling operation that calculates the maximum value for patches of a feature map, and uses it to create a downsampled (pooled) feature map. It is usually…
Average Pooling 설명 없음
Global Average Pooling Global Average Pooling is a pooling operation designed to replace fully connected layers in classical CNNs. The idea is to generate one feature map for each corresponding…
Convolution A convolution is a type of matrix operation, consisting of a kernel, a small matrix of weights, that slides over input data performing element-wise multiplication with the…

Similar Papers 제목 키워드 기반

Divergence of Empirical Neural Tangent Kernel in Classification Problems

2025-04-15 · Zixiong Yu, Songtao Tian, Guhan Chen

This paper demonstrates that in classification problems, fully connected neural networks (FCNs) and residual neural networks (ResNets) cannot be approximated by kernel logistic regression based on the Neural Tangent Kern…

Neural Collapse is Globally Optimal in Deep Regularized ResNets and Transformers

2025-05-21 · Peter Súkeník, Christoph H. Lampert, Marco Mondelli

The empirical emergence of neural collapse -- a surprising symmetry in the feature representations of the training data in the penultimate layer of deep neural networks -- has spurred a line of theoretical research aimed…

Stop Regressing: Training Value Functions via Classification for Scalable Deep RL

2024-03-06 · Jesse Farebrother, Jordi Orbay, Quan Vuong, Adrien Ali Taïga 외

Value functions are a central component of deep reinforcement learning (RL). These functions, parameterized by neural networks, are trained using a mean squared error regression objective to match bootstrapped target val…

Atari GamesDeep Reinforcement LearningregressionReinforcement Learning (RL)

AFFACT - Alignment-Free Facial Attribute Classification Technique

2016-11-18 · Manuel Günther, Andras Rozsa, Terrance E. Boult

Facial attributes are soft-biometrics that allow limiting the search space, e.g., by rejecting identities with non-matching facial characteristics such as nose sizes or eyebrow shapes. In this paper, we investigate how t…

AttributeClassificationData AugmentationFacial Attribute Classification+1

Budgeted Broadcast: An Activity-Dependent Pruning Rule for Neural Network Efficiency

2025-09-26 · Yaron Meirovitch, Fuming Yang, Jeff Lichtman, Nir Shavit arxiv

Most pruning methods remove parameters ranked by impact on loss (e.g., magnitude or gradient). We propose Budgeted Broadcast (BB), which gives each unit a local traffic budget (the product of its long-term on-rate $a_i$ …

Face Identification