paper-with-me

홈 › Papers

How to Score Experts for One-Shot MoE Expert Pruning: A Unified Formulation and Selection Principle

2026-06-14 · Zongfang Liu, Jinghui Zhang, Zijian Ma, Guangyi Chen, Xin Yuan arxiv

Mixture-of-Experts (MoE) language models reduce per-token computation through sparse expert activation, yet deployment still requires storing the full expert pool, making one-shot expert pruning a practical approach for reducing memory usage. Although effective, existing criteria are largely heuristic, and no single criterion is universally optimal. Thus, establishing a principle for selecting pruning criteria suited to different deployment objectives remains an important yet largely underexplored problem in one-shot expert pruning. To this end, we introduce a unified formulation for one-shot MoE expert pruning organized around three factors: routing frequency, gate weighting, and activation strength. The formulation yields a criteria selection principle: task-agnostic pruning should favor routed-token-averaged, gate-free activation-based criteria, whereas task-specific pruning can benefit from retaining routing-frequency and gate-weight information. Beyond this principle, the formulation also provides a systematic view of existing heuristic criteria and gives rise to two new task-agnostic criteria, Mean Activation Norm (MAN) and Mean Squared Activation Norm (MSAN). Across four representative MoE models and 16 diverse benchmarks, MAN and MSAN are consistently strong in the task-agnostic setting, obtain the top-two average ranks, and improve average performance by up to 8.8 points over the strongest baseline.

📄 PDF Abstract BibTeX arXiv:2606.15716

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Generic Expert Coverage for Pruning SparseMixture-of-Experts Language Models

2026-07-02 · Yongqin Zeng, Sicheng Pan, Jiale Wang, Hai-tao Zheng 외 arxiv

Sparsely activated Mixture-of-Experts (MoE) language models contain substantial structured redundancy among routed experts, but pruning them without downstream calibration data remains challenging. Existing expert-prunin…

Finding Fantastic Experts in MoEs: A Unified Study for Expert Dropping Strategies and Observations

2025-04-08 · Ajay Jaiswal, Jianyu Wang, Yixiao Li, Pingzhi Li 외

Sparsely activated Mixture-of-Experts (SMoE) has shown promise in scaling up the learning capacity of neural networks. However, vanilla SMoEs have issues such as expert redundancy and heavy memory requirements, making th…

Instruction FollowingMixture-of-Experts

REAP the Experts: Why Pruning Prevails for One-Shot MoE compression

2025-10-15 · Mike Lasby, Ivan Lazarevich, Nish Sinnadurai, Sean Lie 외 arxiv

Sparsely-activated Mixture-of-Experts (SMoE) models offer efficient pre-training and low latency but their large parameter counts create significant memory overhead, motivating research into expert compression. Contrary …

Code Generation

Cluster-Driven Expert Pruning for Mixture-of-Experts Large Language Models

2025-04-10 · Hongcheng Guo, Juntao Yao, Boyang Wang, Junjia Du 외

Mixture-of-Experts (MoE) architectures have emerged as a promising paradigm for scaling large language models (LLMs) with sparse activation of task-specific experts. Despite their computational efficiency during inferenc…

Computational EfficiencyMixture-of-Experts

Attribution-Guided and Coverage-Maximized Pruning for Structural MoE Compression

2026-06-16 · Yifu Ding, Jiacheng Wang, Ge Yang, Yongcheng Jing 외 arxiv

Mixture-of-Experts (MoE) models scale compute efficiently, yet remain expensive to deploy due to their substantial memory footprint and inference overhead. Prior compression methods mainly operate at the expert level, ei…