paper-with-me

홈 › Papers

LegoNet: Memory Footprint Reduction Through Block Weight Clustering

2026-02-18 · Joseph Bingham, Noah Green, Saman Zonouz arxiv

As the need for neural network-based applications to become more accurate and powerful grows, so too does their size and memory footprint. With embedded devices, whose cache and RAM are limited, this growth hinders their ability to leverage state-of-the-art neural network architectures. In this work, we propose \textbf{LegoNet}, a compression technique that \textbf{constructs blocks of weights of the entire model regardless of layer type} and clusters these induced blocks. Using blocks instead of individual values to cluster the weights, we were able to compress ResNet-50 trained for Cifar-10 and ImageNet with only 32 4x4 blocks, compressing the memory footprint by over a factor of \textbf{64x without having to remove any weights} or changing the architecture and \textbf{no loss to accuracy}, nor retraining or any data, and show how to find an arrangement of 16 4x4 blocks that gives a compression ratio of \textbf{128x with less than 3\% accuracy loss}. This was all achieved with \textbf{no need for (re)training or fine-tuning}.

📄 PDF Abstract BibTeX arXiv:2603.06606

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Structurally Different Neural Network Blocks for the Segmentation of Atrial and Aortic Perivascular Adipose Tissue in Multi-centre CT Angiography Scans

2023-06-06 · Ikboljon Sobirov, Cheng Xie, Muhammad Siddique, Parijat Patel 외

Since the emergence of convolutional neural networks (CNNs) and, later, vision transformers (ViTs), deep learning architectures have predominantly relied on identical block types with varying hyperparameters. We propose …

DiagnosticImage SegmentationMedical Image Segmentationmodel+3

Gefen: Optimized Stochastic Optimizer

2026-06-11 · Nadav Benedek, Tomer Koren, Ohad Fried arxiv

AdamW is a default optimizer for modern deep learning, but its first and second moment states add roughly two parameter-sized buffers to training memory, increasing the already substantial cost of large-scale pretraining…

Efficient In-Memory Acceleration of Sparse Block Diagonal LLMs

2025-10-13 · João Paulo Cardoso de Lima, Marc Dietrich, Jeronimo Castrillon, Asif Ali Khan arxiv

Structured sparsity enables deploying large language models (LLMs) on resource-constrained systems. Approaches like dense-to-sparse fine-tuning are particularly compelling, achieving remarkable structured sparsity by red…

Computational Efficiency

Adam Accumulation to Reduce Memory Footprints of both Activations and Gradients for Large-scale DNN Training

2023-05-31 · Yijia Zhang, Yibo Han, Shijie Cao, Guohao Dai 외

Running out of GPU memory has become a main bottleneck for large-scale DNN training. How to reduce the memory footprint during training has received intensive research attention. We find that previous gradient accumulati…

GPU

BLaST: High Performance Inference and Pretraining using BLock Sparse Transformers

2025-07-03 · Patrik Okanovic, Sameer Deshmukh, Grzegorz Kwasniewski, Yi Zhu 외 arxiv

The energy consumption of large-scale ML models is dominated by data movement, shuffling billions of parameters across memory hierarchies and data centers. Sparsification offers a principled way to mitigate these costs b…