paper-with-me

Papers

Multi-Objective Structured Pruning of LLMs for Latency and Model Size Optimization

2026-06-08 · Muhammad Junaid Ali, Smail Niar, El-Ghazali Talbi arxiv

Large Language Models (LLMs) have achieved widespread adoption because of their strong reasoning and query-response capabilities. However, deploying them in embedded and edge computing environments remains challenging because of strict latency, memory, and energy constraints. Their large parameter counts and computational demands hinder efficient execution on resource-constrained platforms. Although model pruning has emerged as a viable solution for reducing scale while preserving performance, jointly optimizing layers, attention heads, and Multi-Layer Perceptron (MLP) dimensions remains highly complex. Exhaustively exploring this combined design space is computationally expensive and often leads to local optima or unstable configurations. To address these limitations, we propose a hardware-aware, multi-objective structured pruning framework. The proposed two-stage method explicitly targets latency and model size for efficient deployment on edge devices. In the coarse-grained stage, multi-objective depth pruning removes entire attention and MLP blocks to reduce computational load and memory usage. In the subsequent fine-grained stage, Parallel Bayesian Optimization (PBO) searches for the optimal layer-wise pruning ratios for pruning under latency constraints, while importance-based strategies rank the specific components to be pruned within each layer's allocated budget. Experimental results show that our approach reduces model complexity with minimal impact on commonsense reasoning tasks and zero-shot performance. Our method achieves a favorable trade-off among accuracy, latency, and model size, making it suitable for edge deployment. Across multiple LLMs at 37.5% and 50% pruning ratios, the proposed approach achieves better performance on commonsense reasoning tasks than existing methods while significantly reducing inference cost.

📄 PDF Abstract BibTeX arXiv:2607.22583

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Automatic Attention Pruning: Improving and Automating Model Pruning using Attentions

2023-03-14 · Kaiqi Zhao, Animesh Jain, Ming Zhao

Pruning is a promising approach to compress deep learning models in order to deploy them on resource-constrained edge devices. However, many existing pruning solutions are based on unstructured pruning, which yields mode…

Iterative Structured Pruning for Large Language Models with Multi-Domain Calibration

2026-01-06 · Guangxin Wu, Hao Zhang, Zhang Zhibin, Jiafeng Guo 외 arxiv

Large Language Models (LLMs) have achieved remarkable success across a wide spectrum of natural language processing tasks. However, their ever-growing scale introduces significant barriers to real-world deployment, inclu…

Layer-adaptive Structured Pruning Guided by Latency

2023-05-23 · Siyuan Pan, Linna Zhang, Jie Zhang, Xiaoshuang Li 외

Structured pruning can simplify network architecture and improve inference speed. Combined with the underlying hardware and inference engine in which the final model is deployed, better results can be obtained by using l…

Network Pruning

CFSP: An Efficient Structured Pruning Framework for LLMs with Coarse-to-Fine Activation Information

2024-09-20 · Yuxin Wang, Minghua Ma, Zekun Wang, Jingchang Chen 외

The colossal parameters and computational overhead of Large Language Models (LLMs) challenge their real-world applications. Network pruning, which targets unstructured or structured sparsity by removing redundant paramet…

Network Pruning

Dependency-Aware Semi-Structured Sparsity of GLU Variants in Large Language Models

2024-05-03 · Zhiyu Guo, Hidetaka Kamigaito, Taro Wanatnabe

The rapid advancement in Large Language Models (LLMs) has markedly enhanced the capabilities of language understanding and generation. However, the substantial model size poses hardware challenges, affecting both memory …

Computational EfficiencyModel CompressionNetwork Pruning