paper-with-me

홈 › Papers

A Comprehensive and Modularized Statistical Framework for Gradient Norm Equality in Deep Neural Networks

2020-01-01 · Zhaodong Chen, Lei Deng, Bangyan Wang, Guoqi Li, Yuan Xie

In recent years, plenty of metrics have been proposed to identify networks that are free of gradient explosion and vanishing. However, due to the diversity of network components and complex serial-parallel hybrid connections in modern DNNs, the evaluation of existing metrics usually requires strong assumptions, complex statistical analysis, or has limited application fields, which constraints their spread in the community. In this paper, inspired by the Gradient Norm Equality and dynamical isometry, we first propose a novel metric called Block Dynamical Isometry, which measures the change of gradient norm in individual block. Because our Block Dynamical Isometry is norm-based, its evaluation needs weaker assumptions compared with the original dynamical isometry. To mitigate the challenging derivation, we propose a highly modularized statistical framework based on free probability. Our framework includes several key theorems to handle complex serial-parallel hybrid connections and a library to cover the diversity of network components. Besides, several sufficient prerequisites are provided. Powered by our metric and framework, we analyze extensive initialization, normalization, and network structures. We find that Gradient Norm Equality is a universal philosophy behind them. Then, we improve some existing methods based on our analysis, including an activation function selection strategy for initialization techniques, a new configuration for weight normalization, and a depth-aware way to derive coefficients in SeLU. Moreover, we propose a novel normalization technique named second moment normalization, which is theoretically 30% faster than batch normalization without accuracy loss. Last but not least, our conclusions and methods are evidenced by extensive experiments on multiple models over CIFAR10 and ImageNet.

📄 PDF Abstract BibTeX arXiv:2001.00254

Code (1)

apuaaChen/GNEDNN_release 공식 구현 pytorch

Tasks

DiversityPhilosophy

Methods 이 논문이 사용한 방법론

Batch Normalization 설명 없음

Similar Papers 제목 키워드 기반

Multi-Granularity Modularized Network for Abstract Visual Reasoning

2020-07-09 · Xiangru Tang, Haoyuan Wang, Xiang Pan, Jiyang Qi

Abstract visual reasoning connects mental abilities to the physical world, which is a crucial factor in cognitive development. Most toddlers display sensitivity to this skill, but it is not easy for machines. Aimed at it…

Visual GroundingVisual Reasoning

aw_nas: A Modularized and Extensible NAS framework

2020-11-25 · Xuefei Ning, Changcheng Tang, Wenshuo Li, Songyi Yang 외

Neural Architecture Search (NAS) has received extensive attention due to its capability to discover neural network architectures in an automated manner. aw_nas is an open-source Python framework implementing various NAS …

Adversarial RobustnessNeural Architecture Search

CodeChain: Towards Modular Code Generation Through Chain of Self-revisions with Representative Sub-modules

2023-10-13 · Hung Le, Hailin Chen, Amrita Saha, Akash Gokul 외

Large Language Models (LLMs) have already become quite proficient at solving simpler programming tasks like those in HumanEval or MBPP benchmarks. However, solving more complex and competitive programming tasks is still …

Code GenerationHumanEvalmbpp

Improving Computational Complexity in Statistical Models with Second-Order Information

2022-02-09 · Tongzheng Ren, Jiacheng Zhuo, Sujay Sanghavi, Nhat Ho

It is known that when the statistical models are singular, i.e., the Fisher information matrix at the true parameter is degenerate, the fixed step-size gradient descent algorithm takes polynomial number of steps in terms…

parameter estimation

CHULA TTS: A Modularized Text-To-Speech Framework

2014-12-01 · PACLIC 2014 12 · Natthawut Kertkeidkachorn, Supadaech Chanjaradwichai, Proadpran Punyabukkana, Atiwong Suchato
text-to-speechText to Speech