paper-with-me

홈 › Papers

Learning More Generalized Experts by Merging Experts in Mixture-of-Experts

2024-05-19 · Sejik Park

We observe that incorporating a shared layer in a mixture-of-experts can lead to performance degradation. This leads us to hypothesize that learning shared features poses challenges in deep learning, potentially caused by the same feature being learned as various different features. To address this issue, we track each expert's usage frequency and merge the two most frequently selected experts. We then update the least frequently selected expert using the combination of experts. This approach, combined with the subsequent learning of the router's expert selection, allows the model to determine if the most frequently selected experts have learned the same feature differently. If they have, the combined expert can be further trained to learn a more general feature. Consequently, our algorithm enhances transfer learning and mitigates catastrophic forgetting when applied to multi-domain task incremental learning.

📄 PDF Abstract BibTeX arXiv:2405.11530

Code (0)

등록된 구현이 없습니다.

Tasks

Incremental LearningMixture-of-ExpertsTransfer Learning

Similar Papers 제목 키워드 기반

Expert Merging in Sparse Mixture of Experts with Nash Bargaining

2025-10-17 · Dung V. Nguyen, Anh T. Nguyen, Minh H. Nguyen, Luc Q. Nguyen 외 arxiv

Existing expert merging strategies for Sparse Mixture of Experts (SMoE) typically rely on input-dependent or input-independent averaging of expert parameters, but often lack a principled weighting mechanism. In this work…

Image ClassificationText ClassificationLanguage Modelling

Local Mixtures of Experts: Essentially Free Test-Time Training via Model Merging

2025-05-20 · Ryo Bertolissi, Jonas Hübotter, Ido Hakimi, Andreas Krause

Mixture of expert (MoE) models are a promising approach to increasing model capacity without increasing inference cost, and are core components of many state-of-the-art language models. However, current MoE models typica…

Merging Experts into One: Improving Computational Efficiency of Mixture of Experts

2023-10-15 · Shwai He, Run-Ze Fan, Liang Ding, Li Shen 외

Scaling the size of language models usually leads to remarkable advancements in NLP tasks. But it often comes with a price of growing computational cost. Although a sparse Mixture of Experts (MoE) can reduce the cost by …

Computational EfficiencyMixture-of-Experts

Dropping Experts, Recombining Neurons: Retraining-Free Pruning for Sparse Mixture-of-Experts LLMs

2025-09-12 · Yixiao Zhou, Ziyu Zhao, Dongzhou Cheng, zhiliang wu 외 arxiv

Sparse Mixture-of-Experts (SMoE) architectures are widely used in large language models (LLMs) due to their computational efficiency. However, though only a few experts are activated for each token, SMoE still requires l…

Computational Efficiency

Sub-MoE: Efficient Mixture-of-Expert LLMs Compression via Subspace Expert Merging

2025-06-29 · Lujun Li, Zhu Qiyuan, Jiacheng Wang, Wei Li 외

Mixture of Experts (MoE) LLMs face significant obstacles due to their massive parameter scale, which imposes memory, storage, and deployment challenges. Although recent expert merging methods promise greater efficiency b…

Inference OptimizationMixture-of-Experts