paper-with-me

홈 › Papers

MoNE: Replacing Redundant Experts with Lightweight Novices for Structured Pruning of MoE

2025-07-01 · Geng Zhang, Yuxuan Han, Yuxuan Lou, Yiqi Zhang, Wangbo Zhao, Yang You arxiv

Mixture-of-Experts (MoE) enables efficient scaling of large language models by activating only a subset of experts per input token. However, deploying MoE-based models incurs significant memory overhead due to the need to retain all experts in memory. While structured pruning is promising to reduce memory costs, existing methods often show suboptimal performance and unstable degradation in three dimensions: model architectures, calibration data sources, and calibration sample sizes. This paper proposes Mixture-of-Novices-and-Experts (MoNE), a novel expert pruning method that replaces redundant experts with lightweight novices to achieve effective and robust model compression. MoNE evaluates expert redundancy based on two metrics: access frequency and output variance. Experts exhibiting low usage and stable outputs are pruned and replaced with lightweight novices-unbiased estimations of their original outputs-minimizing performance degradation. Extensive experiments demonstrate that MoNE consistently outperforms baseline methods with minimal accuracy degradation across the three dimensions, confirming its effectiveness and robustness. Notably, it outperforms baselines by up to 2.72 for the average zero shot accuracy across nine downstream tasks under 25% pruning ratio, with only 0.14 performance drop for Qwen2-57B-A14B. The code is available at https://github.com/zxgx/mode-pd.

📄 PDF Abstract BibTeX arXiv:2507.00390

Code (0)

등록된 구현이 없습니다.

Tasks

Model Compression

Similar Papers 제목 키워드 기반

Joey NMT: A Minimalist NMT Toolkit for Novices

2019-07-29 · IJCNLP 2019 11 · Julia Kreutzer, Jasmijn Bastings, Stefan Riezler

We present Joey NMT, a minimalist neural machine translation toolkit based on PyTorch that is specifically designed for novices. Joey NMT provides many popular NMT features in a small and simple code base, so that novice…

General KnowledgeMachine TranslationNMTTranslation

Mixture of Nested Experts: Adaptive Processing of Visual Tokens

2024-07-29 · Gagan Jain, Nidhi Hegde, Aditya Kusupati, Arsha Nagrani 외

The visual medium (images and videos) naturally contains a large amount of information redundancy, thereby providing a great opportunity for leveraging efficiency in processing. While Vision Transformer (ViT) based model…

Mixture-of-Experts

Transferring Domain Knowledge with (X)AI-Based Learning Systems

2024-06-03 · Philipp Spitzer, Niklas Kühl, Marc Goutier, Manuel Kaschura 외

In numerous high-stakes domains, training novices via conventional learning systems does not suffice. To impart tacit knowledge, experts' hands-on guidance is imperative. However, training novices by experts is costly an…

Explainable artificial intelligenceExplainable Artificial Intelligence (XAI)

Explaining AI Without Code: A User Study on Explainable AI

2025-12-28 · Natalia Abarca, Andrés Carvallo, Claudia López Moncada, Felipe Bravo-Marquez arxiv

The increasing use of Machine Learning (ML) in sensitive domains such as healthcare, finance, and public policy has raised concerns about the transparency of automated decisions. Explainable AI (XAI) addresses this by cl…

Feature Importance

Training Novices: The Role of Human-AI Collaboration and Knowledge Transfer

2022-07-01 · Philipp Spitzer, Niklas Kühl, Marc Goutier

Across a multitude of work environments, expert knowledge is imperative for humans to conduct tasks with high performance and ensure business success. These humans possess task-specific expert knowledge (TSEK) and hence,…

Transfer Learning