paper-with-me

홈 › Papers

Enhanced Sparsification via Stimulative Training

2024-03-11 · Shengji Tang, Weihao Lin, Hancheng Ye, Peng Ye, Chong Yu, Baopu Li, Tao Chen

Sparsification-based pruning has been an important category in model compression. Existing methods commonly set sparsity-inducing penalty terms to suppress the importance of dropped weights, which is regarded as the suppressed sparsification paradigm. However, this paradigm inactivates the dropped parts of networks causing capacity damage before pruning, thereby leading to performance degradation. To alleviate this issue, we first study and reveal the relative sparsity effect in emerging stimulative training and then propose a structured pruning framework, named STP, based on an enhanced sparsification paradigm which maintains the magnitude of dropped weights and enhances the expressivity of kept weights by self-distillation. Besides, to find an optimal architecture for the pruned network, we propose a multi-dimension architecture space and a knowledge distillation-guided exploration strategy. To reduce the huge capacity gap of distillation, we propose a subnet mutating expansion technique. Extensive experiments on various benchmarks indicate the effectiveness of STP. Specifically, without fine-tuning, our method consistently achieves superior performance at different budgets, especially under extremely aggressive pruning scenarios, e.g., remaining 95.11% Top-1 accuracy (72.43% in 76.15%) while reducing 85% FLOPs for ResNet-50 on ImageNet. Codes will be released soon.

📄 PDF Abstract BibTeX arXiv:2403.06417

Code (0)

등록된 구현이 없습니다.

Tasks

Knowledge DistillationModel Compression

Methods 이 논문이 사용한 방법론

SET Dynamic Sparse Training method where weight mask is updated randomly periodically
Pruning 설명 없음

Similar Papers 제목 키워드 기반

Stimulative Training++: Go Beyond The Performance Limits of Residual Networks

2023-05-04 · Peng Ye, Tong He, Shengji Tang, Baopu Li 외

Residual networks have shown great success and become indispensable in recent deep neural network models. In this work, we aim to re-investigate the training process of residual networks from a novel social psychology pe…

Stimulative Training of Residual Networks: A Social Psychology Perspective of Loafing

2022-10-09 · Peng Ye, Shengji Tang, Baopu Li, Tao Chen 외

Residual networks have shown great success and become indispensable in today's deep models. In this work, we aim to re-investigate the training process of residual networks from a novel social psychology perspective of l…

Exploring the stimulative effect on following drivers in a consecutive lane-change using microscopic vehicle trajectory data

2022-05-18 · Ruifeng Gu

Improper lane-changing behaviors may result in breakdown of traffic flow and the occurrence of various types of collisions. This study investigates lane-changing behaviors of multiple vehicles and the stimulative effect …

Graph Sparsification for Enhanced Conformal Prediction in Graph Neural Networks

2024-10-28 · Yuntian He, Pranav Maneriker, Anutam Srinivasan, Aditya T. Vadlamani 외

Conformal Prediction is a robust framework that ensures reliable coverage across machine learning tasks. Although recent studies have applied conformal prediction to graph neural networks, they have largely emphasized po…

Conformal PredictionDenoisingPrediction

Principle of Relevant Information for Graph Sparsification

2022-05-31 · Shujian Yu, Francesco Alesiani, Wenzhe Yin, Robert Jenssen 외

Graph sparsification aims to reduce the number of edges of a graph while maintaining its structural properties. In this paper, we propose the first general and effective information-theoretic formulation of graph sparsif…

Multi-Task Learning