paper-with-me

Papers

Denoised Variance-Based Pruning with Optimal Brain Bias Compensation

2026-08-18 · Geon Tack Lee, Jaegul Choo, Kang Eun Jeon arxiv

Vision Transformers (ViTs) achieve state-of-the-art performance but carry massive computational overhead that restricts edge deployment. Although structural pruning has emerged as a key strategy to reduce these costs, existing methods often suffer from severe accuracy degradation or require expensive retraining. Recently, Variance-Based Pruning (VBP) introduced a promising paradigm by selecting neurons based on activation variance; however, it remains limited by statistical noise in finite-sample activation covariance and reliance on bias-only updates that cannot fully account for structural reconstruction error. To address these limitations, we introduce Denoised Variance-Based Pruning with Optimal Brain Bias Compensation (DVBP + OB$^2$C). We leverage random matrix theory to filter noise from the activation covariance spectrum for robust neuron selection and mathematically prove that integrating mean-shift compensation into the Optimal Brain Compression objective reduces the layer-wise Hessian exactly to the activation covariance matrix. This enables an optimal, closed-form update of the remaining weights using the same statistics gathered for selection. Extensive experiments on DeiT, Swin, and ConvNeXt architectures demonstrate that DVBP + OB$^2$C achieves state-of-the-art training-free performance; at 50% MLP pruning, it retains over 90% of the original Top-1 accuracy on Small and Base variants, outperforming VBP by up to 29.46% (ConvNeXt-T) and 7.33% (Swin-S). The code is available at: https://github.com/geontackee/DVBP_OB2C.

📄 PDF Abstract BibTeX arXiv:2608.17657

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Pruned non-local means

2017-01-28 · Sanjay Ghosh, Amit K. Mandal, Kunal. N. Chaudhury

In Non-Local Means (NLM), each pixel is denoised by performing a weighted averaging of its neighboring pixels, where the weights are computed using image patches. We demonstrate that the denoising performance of NLM can …

Denoising

Optimal Brain Restoration for Joint Quantization and Sparsification of LLMs

2025-09-14 · Hang Guo, Yawei Li, Luca Benini arxiv

Recent advances in Large Language Model (LLM) compression, such as quantization and pruning, have achieved notable success. However, as these techniques gradually approach their respective limits, relying on a single met…

Minimum Variance Unbiased N:M Sparsity for the Neural Gradients

2022-03-21 · Brian Chmiel, Itay Hubara, Ron Banner, Daniel Soudry

In deep learning, fine-grained N:M sparsity reduces the data footprint and bandwidth of a General Matrix multiply (GEMM) up to x2, and doubles throughput by skipping computation of zero values. So far, it was mainly only…

Optimal Brain Connection: Towards Efficient Structural Pruning

2025-08-07 · Shaowu Chen, Wei Ma, Binhua Huang, Qingyuan Wang 외 arxiv

Structural pruning has been widely studied for its effectiveness in compressing neural networks. However, existing methods often neglect the interconnections among parameters. To address this limitation, this paper propo…

Classification Tree Pruning Under Covariate Shift

2023-05-07 · Nicholas Galbraith, Samory Kpotufe

We consider the problem of \emph{pruning} a classification tree, that is, selecting a suitable subtree that balances bias and variance, in common situations with inhomogeneous training data. Namely, assuming access to mo…

Classification