paper-with-me

홈 › Papers

Robustness in deep learning: The good (width), the bad (depth), and the ugly (initialization)

2022-09-15 · Zhenyu Zhu, Fanghui Liu, Grigorios G Chrysos, Volkan Cevher

We study the average robustness notion in deep neural networks in (selected) wide and narrow, deep and shallow, as well as lazy and non-lazy training settings. We prove that in the under-parameterized setting, width has a negative effect while it improves robustness in the over-parameterized setting. The effect of depth closely depends on the initialization and the training mode. In particular, when initialized with LeCun initialization, depth helps robustness with the lazy training regime. In contrast, when initialized with Neural Tangent Kernel (NTK) and He-initialization, depth hurts the robustness. Moreover, under the non-lazy training regime, we demonstrate how the width of a two-layer ReLU network benefits robustness. Our theoretical developments improve the results by [Huang et al. NeurIPS21; Wu et al. NeurIPS21] and are consistent with [Bubeck and Sellke NeurIPS21; Bubeck et al. COLT21].

📄 PDF Abstract BibTeX arXiv:2209.07263

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

AugLy: Data Augmentations for Robustness

2022-01-17 · Zoe Papakipos, Joanna Bitton

We introduce AugLy, a data augmentation library with a focus on adversarial robustness. AugLy provides a wide array of augmentations for multiple modalities (audio, image, text, & video). These augmentations were inspire…

Adversarial RobustnessData Augmentation

Provable Benefit of Orthogonal Initialization in Optimizing Deep Linear Networks

2020-01-16 · ICLR 2020 1 · Wei Hu, Lechao Xiao, Jeffrey Pennington

The selection of initial parameter values for gradient-based optimization of deep neural networks is one of the most impactful hyperparameter choices in deep learning systems, affecting both convergence times and model p…

The Shaped Transformer: Attention Models in the Infinite Depth-and-Width Limit

2023-06-30 · NeurIPS 2023 11 · Lorenzo Noci, Chuning Li, Mufan Bill Li, Bobby He 외

In deep learning theory, the covariance matrix of the representations serves as a proxy to examine the network's trainability. Motivated by the success of Transformers, we study the covariance matrix of a modified Softma…

Deep AttentionLearning Theory

The Future is Log-Gaussian: ResNets and Their Infinite-Depth-and-Width Limit at Initialization

2021-06-07 · NeurIPS 2021 12 · Mufan Bill Li, Mihai Nica, Daniel M. Roy

Theoretical results show that neural networks can be approximated by Gaussian processes in the infinite-width limit. However, for fully connected networks, it has been previously shown that for any fixed network width, $…

Gaussian Processes

expOSE: Accurate Initialization-Free Projective Factorization Using Exponential Regularization

2023-01-01 · CVPR 2023 1 · José Pedro Iglesias, Amanda Nilsson, Carl Olsson

Bundle adjustment is a key component in practically all available Structure from Motion systems. While it is crucial for achieving accurate reconstruction, convergence to the right solution hinges on good initializat…