paper-with-me

홈 › Papers

On the relationship between topology and gradient propagation in deep networks

2021-01-01 · Kartikeya Bhardwaj, Guihong Li, Radu Marculescu

In this paper, we address two fundamental research questions in neural architecture design: (i) How does the architecture topology impact the gradient flow during training? (ii) Can certain topological characteristics of deep networks indicate a priori (i.e., without training) which models, with a different number of parameters/FLOPS/layers, achieve a similar accuracy? To this end, we formulate the problem of deep learning architecture design from a network science perspective and introduce a new metric called NN-Mass to quantify how effectively information flows through a given architecture. We establish a theoretical link between NN-Mass, a topological property of neural architectures, and gradient flow characteristics (e.g., Layerwise Dynamical Isometry). As such, NN-Mass can identify models with similar accuracy, despite having significantly different size/compute requirements. Detailed experiments on both synthetic and real datasets (e.g., MNIST, CIFAR-10, CIFAR-100, ImageNet) provide extensive evidence for our insights. Finally, we show that the closed-form equation of our theoretically grounded NN-Mass metric enables us to design efficient architectures directly without time-consuming training and searching.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Formal derivation of Mesh Neural Networks with their Forward-Only gradient Propagation

2019-05-16 · Federico A. Galatolo, Mario G. C. A. Cimino, Gigliola Vaglini

This paper proposes the Mesh Neural Network (MNN), a novel architecture which allows neurons to be connected in any topology, to efficiently route information. In MNNs, information is propagated between neurons throughou…

tensor algebra

Scalable Partial Explainability in Neural Networks via Flexible Activation Functions

2020-06-10 · Schyler C. Sun, Chen Li, Zhuangkun Wei, Antonios Tsourdos 외

Achieving transparency in black-box deep learning algorithms is still an open challenge. High dimensional features and decisions given by deep neural networks (NN) require new algorithms and methods to expose its mechani…

Binary ClassificationGaussian Processes

Distributed Training of Graph Convolutional Networks

2020-07-13 · Simone Scardapane, Indro Spinelli, Paolo Di Lorenzo

The aim of this work is to develop a fully-distributed algorithmic framework for training graph convolutional networks (GCNs). The proposed method is able to exploit the meaningful relational structure of the input data,…

Distributed Optimization

Decoupling and Damping: Structurally-Regularized Gradient Matching for Multimodal Graph Condensation

2025-11-25 · Lian Shen, Zhendan Chen, Meijia Song, Yinhui jiang 외 arxiv

In multimodal graph learning, graph structures that integrate information from multiple sources, such as vision and text, can more comprehensively model complex entity relationships. However, the continuous growth of the…

Graph Learning

Leveraging Structural Knowledge in Diffusion Models for Source Localization in Data-Limited Graph Scenarios

2025-02-25 · Hongyi Chen, Jingtao Ding, Xiaojun Liang, Yong Li 외

The source localization problem in graph information propagation is crucial for managing various network disruptions, from misinformation spread to infrastructure failures. While recent deep generative approaches have sh…

DenoisingMisinformation