paper-with-me

홈 › Papers

Averaged Evaluation Masks Capability Trade-Offs: Multi-Source Calibration for High-Sparsity LLM Pruning

2026-06-02 · Hu Xu, Zhaolong Xing, Congcong Liu, Jiaxing Wang, Zhida Jiang, Junshi Huang, Zhen Chen, Jianfeng Xu arxiv

Calibration data are often treated as a minor implementation detail in post-training LLM pruning because averaged evaluations suggest only modest effects. We show that this conclusion is an averaging artifact: at 60\% SparseGPT sparsity, calibration strategies separated by only 2.85 points in averaged commonsense accuracy differ by 51.9 points in Code retention. Across 15 sources, capability-decomposed analysis reveals an opposing pattern: calibration perplexity is positively associated with General retention but negatively associated with Math or Code retention, leaving no evaluated single source uniformly strong across capabilities. This finding motivates capability-balanced multi-source calibration. Under the same calibration budget, a balanced real-data mixture outperforms every evaluated single source on LLaMA-3.1-8B, beating C4 by 18.8 points; the advantage grows with sparsity and persists on LLaMA-3.1-70B. Because the original pretraining data of advanced LLMs are often inaccessible, we further introduce Information-Guided Self-Calibration for Pruning (IGSP). Using only the base model and evaluation taxonomy, IGSP generates capability-stratified pools and selects low-redundancy samples within capability-specific perplexity ranges, outperforming Self-Cal and SGS by up to 4.8 points. Together, these results recast calibration as a capability-coverage problem and identify multi-source design as a practical principle for preserving capabilities in high-sparsity LLM pruning.

📄 PDF Abstract BibTeX arXiv:2606.03328

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Fundamental Limits of Obfuscation for Linear Gaussian Dynamical Systems: An Information-Theoretic Approach

2020-10-29 · Song Fang, Quanyan Zhu

In this paper, we study the fundamental limits of obfuscation in terms of privacy-distortion tradeoffs for linear Gaussian dynamical systems via an information-theoretic approach. Particularly, we obtain analytical formu…

An Information-Theoretic Analysis of Discrete-Time Control and Filtering Limitations by the I-MMSE Relationships

2023-04-18 · Neng Wan, Dapeng Li, Naira Hovakimyan, Petros G. Voulgaris

Fundamental limitations or performance trade-offs/limits are important properties and constraints of both control and filtering systems. Among various trade-off metrics, total information rate that characterizes the sens…

AP-BMM: Approximating Capability-Cost Pareto Sets of LLMs via Asynchronous Prior-Guided Bayesian Model Merging

2025-12-10 · Kesheng Chen, Yamin Hu, Zhenqian Zhu, Yiya Diao 외 arxiv

LLM services need to offer a family of models spanning different capability--cost trade-offs to accommodate diverse user preferences. Model merging offers a practical way to construct such a model family by combining a r…

The influence of the composition of tradeoffs on the generation of differentiated cells

2016-08-30

We study the emergence of cell differentiation under the assumption of the existence of a given number of tradeoffs between genes encoding different functions. In the model the viability of colonies is determined by the …

Personality as a Probe for LLM Evaluation: Method Trade-offs and Downstream Effects

2025-09-05 · Gunmay Handa, Zekun Wu, Adriano Koshiyama, Philip Treleaven arxiv

Personality manipulation in large language models (LLMs) is increasingly applied in customer service and agentic scenarios, yet its mechanisms and trade-offs remain unclear. We present a systematic study of personality c…

parameter-efficient fine-tuning