paper-with-me

홈 › Papers

Dynamic Mixture of Experts: An Auto-Tuning Approach for Efficient Transformer Models

2024-05-23 · Yongxin Guo, Zhenglin Cheng, Xiaoying Tang, Zhaopeng Tu, Tao Lin

The Sparse Mixture of Experts (SMoE) has been widely employed to enhance the efficiency of training and inference for Transformer-based foundational models, yielding promising results. However, the performance of SMoE heavily depends on the choice of hyper-parameters, such as the number of experts and the number of experts to be activated (referred to as top-k), resulting in significant computational overhead due to the extensive model training by searching over various hyper-parameter configurations. As a remedy, we introduce the Dynamic Mixture of Experts (DynMoE) technique. DynMoE incorporates (1) a novel gating method that enables each token to automatically determine the number of experts to activate. (2) An adaptive process automatically adjusts the number of experts during training. Extensive numerical results across Vision, Language, and Vision-Language tasks demonstrate the effectiveness of our approach to achieve competitive performance compared to GMoE for vision and language tasks, and MoE-LLaVA for vision-language tasks, while maintaining efficiency by activating fewer parameters. Our code is available at https://github.com/LINs-lab/DynMoE.

📄 PDF Abstract BibTeX arXiv:2405.14297

Code (1)

lins-lab/dynmoe 공식 구현 pytorch

Tasks

Mixture-of-ExpertsVisual Question Answering

Similar Papers 제목 키워드 기반

DA-MoE: Towards Dynamic Expert Allocation for Mixture-of-Experts Models

2024-09-10 · Maryam Akhavan Aghdam, Hongpeng Jin, Yanzhao Wu

Transformer-based Mixture-of-Experts (MoE) models have been driving several recent technological advancements in Natural Language Processing (NLP). These MoE models adopt a router mechanism to determine which experts to …

Mixture-of-Experts

Dynamic Mixture-of-Experts for Visual Autoregressive Model

2025-10-08 · Jort Vincenti, Metod Jazbec, Guoxuan Xia arxiv

Visual Autoregressive Models (VAR) offer efficient and high-quality image generation but suffer from computational redundancy due to repeated Transformer calls at increasing resolutions. We introduce a dynamic Mixture-of…

Image Generation

MixtureKit: A General Framework for Composing, Training, and Visualizing Mixture-of-Experts Models

2025-12-13 · Ahmad Chamma, Omar El Herraoui, Guokan Shang arxiv

We introduce MixtureKit, a modular open-source framework for constructing, training, and analyzing Mixture-of-Experts (MoE) models from arbitrary pre-trained or fine-tuned models. MixtureKit currently supports three comp…

Mixture-of-Control: State-Aware Fine-Tuning for Transformer-based Models

2026-06-30 · Duc Anh Nguyen, Tien Ngoc Luu, Tung Pham, Toan Tran arxiv

State-based fine-tuning has emerged as a compelling alternative to weight-based adaptation for transformers, updating lightweight controls into states rather than model weights, offering substantial memory savings while …

Computational EfficiencyRepresentation Learning

Efficient Fine-tuning of Audio Spectrogram Transformers via Soft Mixture of Adapters

2024-02-01 · Umberto Cappellazzo, Daniele Falavigna, Alessio Brutti

Mixture of Experts (MoE) architectures have recently started burgeoning due to their ability to scale model's capacity while maintaining the computational cost affordable. Furthermore, they can be applied to both Transfo…

Mixture-of-Expertsparameter-efficient fine-tuningState Space ModelsTransfer Learning