paper-with-me

Papers

Get More at Once: Alternating Sparse Training with Gradient Correction

2022-11-01 · NIPS 2022 11 · Li Yang, Jian Meng, Jae-sun Seo, Deliang Fan

Recently, a new trend of exploring training sparsity has emerged, which remove parameters during training, leading to both training and inference efficiency improvement. This line of works primarily aims to obtain a single sparse model under a pre-defined large sparsity ratio. It leads to a static/fixed sparse inference model that is not capable of adjusting or re-configuring its computation complexity (i.e., inference structure, latency) after training for real-world varying and dynamic hardware resource availability. To enable such run-time or post-training network morphing, the concept of dynamic inference' or training-once-for-all' has been proposed to train a single network consisting of multiple sub-nets once, but each sub-net could perform the same inference function with different computing complexity. However, the traditional dynamic inference training method requires a joint training scheme with multi-objective optimization, which suffers from very large training overhead. In this work, for the first time, we propose a novel alternating sparse training (AST) scheme to train multiple sparse sub-nets for dynamic inference without extra training cost compared to the case of training a single sparse model from scratch. Furthermore, to mitigate the interference of weight update among sub-nets, we propose gradient correction within the inner-group iterations to reduce their weight update interference. We validate the proposed AST on multiple datasets against state-of-the-art sparse training method, which shows that AST achieves similar or better accuracy, but only needs to train once to get multiple sparse sub-nets with different sparsity ratios. More importantly, compared with the traditional joint training based dynamic inference training methodology, the large training overhead is completely eliminated without affecting the accuracy of each sub-net.

📄 PDF Abstract BibTeX

Code (1)

mengjian0502/AST pytorch

Similar Papers 제목 키워드 기반

Clustering with feature selection using alternating minimization, Application to computational biology

2017-11-08 · Cyprien Gilet, Marie Deprez, Jean-Baptiste Caillau, Michel Barlaud

This paper deals with unsupervised clustering with feature selection. The problem is to estimate both labels and a sparse projection matrix of weights. To address this combinatorial non-convex problem maintaining a stric…

Clusteringfeature selection

An Alternating Manifold Proximal Gradient Method for Sparse PCA and Sparse CCA

2019-03-27 · Shixiang Chen, Shiqian Ma, Lingzhou Xue, Hui Zou

Sparse principal component analysis (PCA) and sparse canonical correlation analysis (CCA) are two essential techniques from high-dimensional statistics and machine learning for analyzing large-scale data. Both problems c…

Efficient proximal gradient algorithms for joint graphical lasso

2021-07-16 · Jie Chen, Ryosuke Shimmura, Joe Suzuki

We consider learning an undirected graphical model from sparse data. While several efficient algorithms have been proposed for graphical lasso (GL), the alternating direction method of multipliers (ADMM) is the main appr…

Over-the-Air Federated Learning Over MIMO Channels: A Sparse-Coded Multiplexing Approach

2023-04-10 · Chenxi Zhong, Xiaojun Yuan

The communication bottleneck of over-the-air federated learning (OA-FL) lies in uploading the gradients of local learning models. In this paper, we study the reduction of the communication overhead in the gradients uploa…

Federated Learning

A Scale Invariant Approach for Sparse Signal Recovery

2018-12-20 · Yaghoub Rahimi, Chao Wang, Hongbo Dong, Yifei Lou

In this paper, we study the ratio of the $L_1 $ and $L_2 $ norms, denoted as $L_1/L_2$, to promote sparsity. Due to the non-convexity and non-linearity, there has been little attention to this scale-invariant model. Comp…

MRI Reconstruction