paper-with-me

홈 › Papers

Adaptive Pruning for Large Language Models with Structural Importance Awareness

2024-12-19 · Haotian Zheng, Jinke Ren, Yushan Sun, Ruichen Zhang, Wenbo Zhang, Zhen Li, Dusit Niyato, Shuguang Cui, Yatong Han

The recent advancements in large language models (LLMs) have significantly improved language understanding and generation capabilities. However, it is difficult to deploy LLMs on resource-constrained edge devices due to their high computational and storage resource demands. To address this issue, we propose a novel LLM model pruning method, namely structurally-aware adaptive pruning (SAAP), to significantly reduce the computational and memory costs while maintaining model performance. We first define an adaptive importance fusion metric to evaluate the importance of all coupled structures in LLMs by considering their homoscedastic uncertainty. Then, we rank the importance of all modules to determine the specific layers that should be pruned to meet particular performance requirements. Furthermore, we develop a new group fine-tuning strategy to improve the inference efficiency of LLMs. Finally, we evaluate the proposed SAAP method on multiple LLMs across two common tasks, i.e., zero-shot classification and text generation. Experimental results show that our SAAP method outperforms several state-of-the-art baseline methods, achieving 2.17%, 2.37%, and 2.39% accuracy gains on LLaMA-7B, Vicuna-7B, and LLaMA-13B. Additionally, SAAP improves the token generation speed by 5%, showcasing its practical advantages in resource-constrained scenarios.

📄 PDF Abstract BibTeX arXiv:2412.15127

Code (0)

등록된 구현이 없습니다.

Tasks

Text Generationzero-shot-classificationZero-Shot Learning

Methods 이 논문이 사용한 방법론

SPEED The monocular depth estimation (MDE) is the task of estimating depth from a single frame. This information is an essential knowledge in many computer vision tasks such as scene…
Pruning 설명 없음

Similar Papers 제목 키워드 기반

LLM-BIP: Structured Pruning for Large Language Models with Block-Wise Forward Importance Propagation

2024-12-09 · Haihang Wu

Large language models (LLMs) have demonstrated remarkable performance across various language tasks, but their widespread deployment is impeded by their large size and high computational costs. Structural pruning is a pr…

Sample-aware Adaptive Structured Pruning for Large Language Models

2025-03-08 · Jun Kong, Xinge Ma, Jin Wang, Xuejie Zhang

Large language models (LLMs) have achieved outstanding performance in natural language processing, but enormous model sizes and high computational costs limit their practical deployment. Structured pruning can effectivel…

Bayesian Optimization

Towards Efficient VLMs: Information-Theoretic Driven Compression via Adaptive Structural Pruning

2025-11-24 · Zhaoqi Xu, Yingying Zhang, Jian Li, Jianwei Guo 외 arxiv

Recent advances in vision-language models (VLMs) have shown remarkable performance across multimodal tasks, yet their ever-growing scale poses severe challenges for deployment and efficiency. Existing compression methods…

Toward Adaptive Large Language Models Structured Pruning via Hybrid-grained Weight Importance Assessment

2024-03-16 · Jun Liu, Zhenglun Kong, Pu Zhao, Changdi Yang 외

Structured pruning for large language models (LLMs) has garnered significant academic interest due to its ability to efficiently compress and accelerate LLMs by eliminating redundant weight groups at a coarse-grained gra…

DecoderLanguage ModellingLarge Language Model

Fluctuation-based Adaptive Structured Pruning for Large Language Models

2023-12-19 · Yongqi An, Xu Zhao, Tao Yu, Ming Tang 외

Network Pruning is a promising way to address the huge computing resource demands of the deployment and inference of Large Language Models (LLMs). Retraining-free is important for LLMs' pruning methods. However, almost a…

Network Pruning