paper-with-me

홈 › Papers

The Unreasonable Effectiveness of Random Pruning: Return of the Most Naive Baseline for Sparse Training

2022-02-05 · ICLR 2022 4 · Shiwei Liu, Tianlong Chen, Xiaohan Chen, Li Shen, Decebal Constantin Mocanu, Zhangyang Wang, Mykola Pechenizkiy

Random pruning is arguably the most naive way to attain sparsity in neural networks, but has been deemed uncompetitive by either post-training pruning or sparse training. In this paper, we focus on sparse training and highlight a perhaps counter-intuitive finding, that random pruning at initialization can be quite powerful for the sparse training of modern neural networks. Without any delicate pruning criteria or carefully pursued sparsity structures, we empirically demonstrate that sparsely training a randomly pruned network from scratch can match the performance of its dense equivalent. There are two key factors that contribute to this revival: (i) the network sizes matter: as the original dense networks grow wider and deeper, the performance of training a randomly pruned sparse network will quickly grow to matching that of its dense equivalent, even at high sparsity ratios; (ii) appropriate layer-wise sparsity ratios can be pre-chosen for sparse training, which shows to be another important performance booster. Simple as it looks, a randomly pruned subnetwork of Wide ResNet-50 can be sparsely trained to outperforming a dense Wide ResNet-50, on ImageNet. We also observed such randomly pruned networks outperform dense counterparts in other favorable aspects, such as out-of-distribution detection, uncertainty estimation, and adversarial robustness. Overall, our results strongly suggest there is larger-than-expected room for sparse training at scale, and the benefits of sparsity might be more universal beyond carefully designed pruning. Our source code can be found at https://github.com/VITA-Group/Random_Pruning.

📄 PDF Abstract BibTeX arXiv:2202.02643

Code (1)

vita-group/random_pruning 공식 구현 pytorch

Tasks

Adversarial RobustnessOut-of-Distribution Detection

Methods 이 논문이 사용한 방법론

Pruning 설명 없음

Similar Papers 제목 키워드 기반

The unreasonable effectiveness of pattern matching

2026-01-16 · Gary Lupyan, Blaise Agüera y Arcas arxiv

We report on an astonishing ability of large language models (LLMs) to make sense of "Jabberwocky" language in which most or all content words have been randomly replaced by nonsense strings, e.g., translating "He dwushe…

The unreasonable effectiveness of optimal transport in economics

2021-07-09 · Alfred Galichon

Optimal transport has become part of the standard quantitative economics toolbox. It is the framework of choice to describe models of matching with transfers, but beyond that, it allows to: extend quantile regression; id…

Discrete Choice Modelsquantile regressionregression

The Unreasonable Effectiveness of Random Target Embeddings for Continuous-Output Neural Machine Translation

2023-10-31 · Evgeniia Tokarchuk, Vlad Niculae

Continuous-output neural machine translation (CoNMT) replaces the discrete next-word prediction problem with an embedding prediction. The semantic structure of the target embedding space (i.e., closeness of related words…

Machine TranslationPredictionTranslation

The Unreasonable Ineffectiveness of the Deeper Layers

2024-03-26 · Andrey Gromov, Kushal Tirumala, Hassan Shapourian, Paolo Glorioso 외

We empirically study a simple layer-pruning strategy for popular families of open-weight pretrained LLMs, finding minimal degradation of performance on different question-answering benchmarks until after a large fraction…

GPUQuantizationQuestion Answering

The Unreasonable Effectiveness of Structured Random Orthogonal Embeddings

2017-03-02 · NeurIPS 2017 12 · Krzysztof Choromanski, Mark Rowland, Adrian Weller

We examine a class of embeddings based on structured random matrices with orthogonal rows which can be applied in many machine learning applications including dimensionality reduction and kernel approximation. For both t…

BIG-bench Machine LearningDimensionality Reduction