paper-with-me

Papers

LLMs can Compress LLMs: Adaptive Pruning by Agents

2026-01-14 · Sai Varun Kodathala, Rakesh Vunnam arxiv

As Large Language Models (LLMs) continue to scale, post-training pruning has emerged as a promising approach to reduce computational costs while preserving performance. Existing methods such as SparseGPT and Wanda achieve high sparsity through layer-wise weight reconstruction or activation-aware magnitude pruning, but rely on uniform or hand-crafted heuristics to determine per-layer sparsity ratios. Moreover, recent work has shown that pruned LLMs suffer from severe factual knowledge degradation, with structured pruning methods experiencing near-total collapse in factual question-answering capabilities. We introduce agent-guided pruning, where a foundation model acts as an adaptive pruning agent to intelligently select which layers to prune at each iteration while preserving critical knowledge pathways. Our method constructs layer-wise sensitivity profiles by combining Wanda-inspired weight-activation metrics with gradient importance scores, normalized as z-scores for model-agnostic comparison. These statistics are processed by an LLM agent equipped with self-reflection capabilities, enabling it to learn from previous pruning outcomes and iteratively refine its strategy. A checkpoint rollback mechanism maintains model quality by reverting when perplexity degradation exceeds a threshold. We evaluate our approach on Qwen3 models (4B and 8B parameters) at approximately 45% sparsity, demonstrating substantial improvements over structured pruning baselines: 56% relative improvement in MMLU accuracy, 19x better factual knowledge retention on FreebaseQA, and 69% lower perplexity degradation. Notably, our framework requires no retraining, operates in a model-agnostic manner, and exhibits effective self-correction with only 2-4 rollbacks across 21-40 iterations, demonstrating that foundation models can effectively guide the compression of other foundation models.

📄 PDF Abstract BibTeX arXiv:2601.09694

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Fluctuation-based Adaptive Structured Pruning for Large Language Models

2023-12-19 · Yongqi An, Xu Zhao, Tao Yu, Ming Tang 외

Network Pruning is a promising way to address the huge computing resource demands of the deployment and inference of Large Language Models (LLMs). Retraining-free is important for LLMs' pruning methods. However, almost a…

Network Pruning

Toward Adaptive Large Language Models Structured Pruning via Hybrid-grained Weight Importance Assessment

2024-03-16 · Jun Liu, Zhenglun Kong, Pu Zhao, Changdi Yang 외

Structured pruning for large language models (LLMs) has garnered significant academic interest due to its ability to efficiently compress and accelerate LLMs by eliminating redundant weight groups at a coarse-grained gra…

DecoderLanguage ModellingLarge Language Model

Sharp Eyes and Memory for VideoLLMs: Information-Aware Visual Token Pruning for Efficient and Reliable VideoLLM Reasoning

2025-11-11 · Jialong Qin, Xin Zou, Di Lu, Yibo Yan 외 arxiv

Current Video Large Language Models (VideoLLMs) suffer from quadratic computational complexity and key-value cache scaling, due to their reliance on processing excessive redundant visual tokens. To address this problem, …

One-Shot Sensitivity-Aware Mixed Sparsity Pruning for Large Language Models

2023-10-14 · Hang Shao, Bei Liu, Bo Xiao, Ke Zeng 외

Various Large Language Models~(LLMs) from the Generative Pretrained Transformer(GPT) family have achieved outstanding performances in a wide range of text generation tasks. However, the enormous model sizes have hindered…

QuantizationSensitivityText Generation

OPTISHEAR: Towards Efficient and Adaptive Pruning of Large Language Models via Evolutionary Optimization

2025-02-15 · Shuqi Liu, Bowei He, Han Wu, Linqi Song

Post-training pruning has emerged as a crucial optimization technique as large language models (LLMs) continue to grow rapidly. However, the significant variations in weight distributions across different LLMs make fixed…

Model Compression