paper-with-me

홈 › Papers

Sparsity Winning Twice: Better Robust Generalization from More Efficient Training

2022-02-20 · ICLR 2022 4 · Tianlong Chen, Zhenyu Zhang, Pengjun Wang, Santosh Balachandra, Haoyu Ma, Zehao Wang, Zhangyang Wang

Recent studies demonstrate that deep networks, even robustified by the state-of-the-art adversarial training (AT), still suffer from large robust generalization gaps, in addition to the much more expensive training costs than standard training. In this paper, we investigate this intriguing problem from a new perspective, i.e., injecting appropriate forms of sparsity during adversarial training. We introduce two alternatives for sparse adversarial training: (i) static sparsity, by leveraging recent results from the lottery ticket hypothesis to identify critical sparse subnetworks arising from the early training; (ii) dynamic sparsity, by allowing the sparse subnetwork to adaptively adjust its connectivity pattern (while sticking to the same sparsity ratio) throughout training. We find both static and dynamic sparse methods to yield win-win: substantially shrinking the robust generalization gap and alleviating the robust overfitting, meanwhile significantly saving training and inference FLOPs. Extensive experiments validate our proposals with multiple network architectures on diverse datasets, including CIFAR-10/100 and Tiny-ImageNet. For example, our methods reduce robust generalization gap and overfitting by 34.44% and 4.02%, with comparable robust/standard accuracy boosts and 87.83%/87.82% training/inference FLOPs savings on CIFAR-100 with ResNet-18. Besides, our approaches can be organically combined with existing regularizers, establishing new state-of-the-art results in AT. Codes are available in https://github.com/VITA-Group/Sparsity-Win-Robust-Generalization.

📄 PDF Abstract BibTeX arXiv:2202.09844

Code (1)

vita-group/sparsity-win-robust-generalization 공식 구현 pytorch

Similar Papers 제목 키워드 기반

Advancing Model Pruning via Bi-level Optimization

2022-10-08 · Yihua Zhang, Yuguang Yao, Parikshit Ram, Pu Zhao 외

The deployment constraints in practical applications necessitate the pruning of large-scale deep learning models, i.e., promoting their weight sparsity. As illustrated by the Lottery Ticket Hypothesis (LTH), pruning also…

model

Lottery Tickets can have Structural Sparsity

2021-09-29 · Tianlong Chen, Xuxi Chen, Xiaolong Ma, Yanzhi Wang 외

The lottery ticket hypothesis (LTH) has shown that dense models contain highly sparse subnetworks (i.e., $\textit{winning tickets}$) that can be trained in isolation to match full accuracy. Despite many exciting efforts …

Successfully Applying Lottery Ticket Hypothesis to Diffusion Model

2023-10-28 · Chao Jiang, Bo Hui, Bohan Liu, Da Yan

Despite the success of diffusion models, the training and inference of diffusion models are notoriously expensive due to the long chain of the reverse process. In parallel, the Lottery Ticket Hypothesis (LTH) claims that…

Denoising

Sparse Double Descent: Where Network Pruning Aggravates Overfitting

2022-06-17 · Zheng He, Zeke Xie, Quanzhi Zhu, Zengchang Qin

People usually believe that network pruning not only reduces the computational cost of deep networks, but also prevents overfitting by decreasing model capacity. However, our work surprisingly discovers that network prun…

Network Pruning

Coarsening the Granularity: Towards Structurally Sparse Lottery Tickets

2022-02-09 · Tianlong Chen, Xuxi Chen, Xiaolong Ma, Yanzhi Wang 외

The lottery ticket hypothesis (LTH) has shown that dense models contain highly sparse subnetworks (i.e., winning tickets) that can be trained in isolation to match full accuracy. Despite many exciting efforts being made,…