paper-with-me

Papers

The Asymmetric Maximum Margin Bias of Quasi-Homogeneous Neural Networks

2022-10-07 · Daniel Kunin, Atsushi Yamamura, Chao Ma, Surya Ganguli

In this work, we explore the maximum-margin bias of quasi-homogeneous neural networks trained with gradient flow on an exponential loss and past a point of separability. We introduce the class of quasi-homogeneous models, which is expressive enough to describe nearly all neural networks with homogeneous activations, even those with biases, residual connections, and normalization layers, while structured enough to enable geometric analysis of its gradient dynamics. Using this analysis, we generalize the existing results of maximum-margin bias for homogeneous networks to this richer class of models. We find that gradient flow implicitly favors a subset of the parameters, unlike in the case of a homogeneous model where all parameters are treated equally. We demonstrate through simple examples how this strong favoritism toward minimizing an asymmetric norm can degrade the robustness of quasi-homogeneous models. On the other hand, we conjecture that this norm-minimization discards, when possible, unnecessary higher-order parameters, reducing the model to a sparser parameterization. Lastly, by applying our theorem to sufficiently expressive neural networks with normalization layers, we reveal a universal mechanism behind the empirical phenomenon of Neural Collapse.

📄 PDF Abstract BibTeX arXiv:2210.03820

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Stabilizing spiral structures in the asymmetric May-Leonard model

2021-05-17 · Shannon R. Serrao, Uwe C. Täuber

We study the induction and stabilization of spiral structures for the cyclic three-species stochastic May-Leonard model with asymmetric predation rates on a spatially inhomogeneous two-dimensional toroidal lattice using …

A Wideband Quasi-Asymmetric Doherty Power Amplifier with a Two-Section Matching-Phase Difference Compensator Network Design Using GaAs Technology

2020-05-06 · Seyedehmarzieh Rouhani, Ahmad Ghanaatian, Adib Abrishamifar, Majid Tayarani

In this paper, a quasi-asymmetric Doherty power amplifier (PA) is designed without load modulation using the GaAs 0.25{\mu}m pHEMT technology to reach an enlarged output power back-off (OPBO) with circuitry solutions in …

Flavors of Margin: Implicit Bias of Steepest Descent in Homogeneous Neural Networks

2024-10-29 · Nikolaos Tsilivis, Gal Vardi, Julia Kempe

We study the implicit bias of the general family of steepest descent algorithms with infinitesimal learning rate in deep homogeneous neural networks. We show that: (a) an algorithm-dependent geometric margin starts incre…

The Implicit Bias of Adam and Muon on Smooth Homogeneous Neural Networks

2026-02-18 · Eitan Gronich, Gal Vardi arxiv

We study the implicit bias of momentum-based optimizers on smooth homogeneous models. We show that \textit{momentum steepest descent} algorithms like Muon (spectral norm), MomentumGD ($\ell_2$ norm), and Signum ($\ell_\i…

Implicit Bias of Gradient Descent for Non-Homogeneous Deep Networks

2025-02-22 · Yuhang Cai, Kangjie Zhou, Jingfeng Wu, Song Mei 외

We establish the asymptotic implicit bias of gradient descent (GD) for generic non-homogeneous deep networks under exponential loss. Specifically, we characterize three key properties of GD iterates starting from a suffi…