paper-with-me

홈 › Papers

Post-training deep neural network pruning via layer-wise calibration

2021-04-30 · Ivan Lazarevich, Alexander Kozlov, Nikita Malinin

We present a post-training weight pruning method for deep neural networks that achieves accuracy levels tolerable for the production setting and that is sufficiently fast to be run on commodity hardware such as desktop CPUs or edge devices. We propose a data-free extension of the approach for computer vision models based on automatically-generated synthetic fractal images. We obtain state-of-the-art results for data-free neural network pruning, with ~1.5% top@1 accuracy drop for a ResNet50 on ImageNet at 50% sparsity rate. When using real data, we are able to get a ResNet50 model on ImageNet with 65% sparsity rate in 8-bit precision in a post-training setting with a ~1% top@1 accuracy drop. We release the code as a part of the OpenVINO(TM) Post-Training Optimization tool.

📄 PDF Abstract BibTeX arXiv:2104.15023

Code (0)

등록된 구현이 없습니다.

Tasks

Network Pruning

Methods 이 논문이 사용한 방법론

Pruning 설명 없음

Similar Papers 제목 키워드 기반

Adaptive Signal Resuscitation: Channel-wise Post-Pruning Repair for Sparse Vision Networks

2026-05-20 · Qishi Zhan, Ziheng Chen, Minxuan Hu arxiv

One-shot magnitude pruning can cause severe accuracy collapse in the high-sparsity regime, even when the pruning mask preserves the largest weights. We argue that this failure reflects a granularity mismatch in post-prun…

Relative Repairability: A Calibration-Based Diagnostic for High-Sparsity Post-Pruning Allocation

2026-05-25 · Qishi Zhan, Liang He, Minxuan Hu, Ziheng Chen arxiv

At very high sparsity, neural network pruning does more than decide which weights remain. It also determines where pruning induced damage is placed across the network, and whether that damage can be recovered by a fixed …

Network Pruning

On the Impact of Calibration Data in Post-training Quantization and Pruning

2023-11-16 · Miles Williams, Nikolaos Aletras

Quantization and pruning form the foundation of compression for neural networks, enabling efficient inference for large language models (LLMs). Recently, various quantization and pruning techniques have demonstrated rema…

Model CompressionQuantization

From Local to Global: Revisiting Structured Pruning Paradigms for Large Language Models

2025-10-20 · Ziyan Wang, Enmao Diao, Qi Le, Pu Wang 외 arxiv

Structured pruning is a practical approach to deploying large language models (LLMs) efficiently, as it yields compact, hardware-friendly architectures. However, the dominant local paradigm is task-agnostic: by optimizin…

UniPruning: Unifying Local Metric and Global Feedback for Scalable Sparse LLMs

2025-09-29 · Yizhuo Ding, Wanying Qu, Jiawei Geng, Wenqi Shao 외 arxiv

Large Language Models (LLMs) achieve strong performance across diverse tasks but face prohibitive computational and memory costs. Pruning offers a promising path by inducing sparsity while preserving architectural flexib…