paper-with-me

홈 › Papers

A Novel Structure-Agnostic Multi-Objective Approach for Weight-Sharing Compression in Deep Neural Networks

2025-01-06 · Rasa Khosrowshahli, Shahryar Rahnamayan, Beatrice Ombuki-Berman

Deep neural networks suffer from storing millions and billions of weights in memory post-training, making challenging memory-intensive models to deploy on embedded devices. The weight-sharing technique is one of the popular compression approaches that use fewer weight values and share across specific connections in the network. In this paper, we propose a multi-objective evolutionary algorithm (MOEA) based compression framework independent of neural network architecture, dimension, task, and dataset. We use uniformly sized bins to quantize network weights into a single codebook (lookup table) for efficient weight representation. Using MOEA, we search for Pareto optimal $k$ bins by optimizing two objectives. Then, we apply the iterative merge technique to non-dominated Pareto frontier solutions by combining neighboring bins without degrading performance to decrease the number of bins and increase the compression ratio. Our approach is model- and layer-independent, meaning the weights are mixed in the clusters from any layer, and the uniform quantization method used in this work has $O(N)$ complexity instead of non-uniform quantization methods such as k-means with $O(Nkt)$ complexity. In addition, we use the center of clusters as the shared weight values instead of retraining shared weights, which is computationally expensive. The advantage of using evolutionary multi-objective optimization is that it can obtain non-dominated Pareto frontier solutions with respect to performance and shared weights. The experimental results show that we can reduce the neural network memory by $13.72 \sim14.98 \times$ on CIFAR-10, $11.61 \sim 12.99\times$ on CIFAR-100, and $7.44 \sim 8.58\times$ on ImageNet showcasing the effectiveness of the proposed deep neural network compression framework.

📄 PDF Abstract BibTeX arXiv:2501.03095

Code (0)

등록된 구현이 없습니다.

Tasks

Neural Network CompressionQuantization

Similar Papers 제목 키워드 기반

On Weight-Sharing and Bilevel Optimization in Architecture Search

2019-09-25 · Mikhail Khodak, Liam Li, Maria-Florina Balcan, Ameet Talwalkar

Weight-sharing—the simultaneous optimization of multiple neural networks using the same parameters—has emerged as a key component of state-of-the-art neural architecture search. However, its success is poorly understood …

Bilevel Optimizationfeature selectionNeural Architecture Search

AutoDistil: Few-shot Task-agnostic Neural Architecture Search for Distilling Large Language Models

2022-01-29 · Dongkuan Xu, Subhabrata Mukherjee, Xiaodong Liu, Debadeepta Dey 외

Knowledge distillation (KD) methods compress large models into smaller students with manually-designed student architectures given pre-specified computational cost. This requires several trials to find a viable student, …

Inductive BiasKnowledge DistillationNeural Architecture Search

How Does Supernet Help in Neural Architecture Search?

2020-10-16 · Yuge Zhang, Quanlu Zhang, Yaming Yang

Weight sharing, as an approach to speed up architecture performance estimation has received wide attention. Instead of training each architecture separately, weight sharing builds a supernet that assembles all the archit…

Neural Architecture Search

Mixture of Universal Experts: Scaling Virtual Width via Depth-Width Transformation

2026-03-05 · Yilong Chen, Naibin Gu, Junyuan Shang, Zhenyu Zhang 외 arxiv

Mixture-of-Experts (MoE) decouples model capacity from per-token computation, yet their scalability remains limited by the physical dimensions of depth and width. To overcome this, we propose Mixture of Universal Experts…

Enhancing Explicit and Implicit Feature Interactions via Information Sharing for Parallel Deep CTR Models

2021-10-30 · Proceedings of the 30th ACM International Conference on Information & Knowledge Management 2021 10 · Bo Chen, Yichao Wang, Zhirong Liu, Ruiming Tang 외

Effectively modeling feature interactions is crucial for CTR prediction in industrial recommender systems. The state-of-the-art deep CTR models with parallel structure (e.g., DCN) learn explicit and implicit feature inte…

Click-Through Rate PredictionRecommendation Systems