paper-with-me

홈 › Papers

Ada-K Routing: Boosting the Efficiency of MoE-based LLMs

2024-10-14 · Tongtian Yue, Longteng Guo, Jie Cheng, Xuange Gao, Jing Liu

In the era of Large Language Models (LLMs), Mixture-of-Experts (MoE) architectures offer a promising approach to managing computational costs while scaling up model parameters. Conventional MoE-based LLMs typically employ static Top-K routing, which activates a fixed and equal number of experts for each token regardless of their significance within the context. In this paper, we propose a novel Ada-K routing strategy that dynamically adjusts the number of activated experts for each token, thereby improving the balance between computational efficiency and model performance. Specifically, our strategy incorporates learnable and lightweight allocator modules that decide customized expert resource allocation tailored to the contextual needs for each token. These allocators are designed to be fully pluggable, making it broadly applicable across all mainstream MoE-based LLMs. We leverage the Proximal Policy Optimization (PPO) algorithm to facilitate an end-to-end learning process for this non-differentiable decision-making framework. Extensive evaluations on four popular baseline models demonstrate that our Ada-K routing method significantly outperforms conventional Top-K routing. Compared to Top-K, our method achieves over 25% reduction in FLOPs and more than 20% inference speedup while still improving performance across various benchmarks. Moreover, the training of Ada-K is highly efficient. Even for Mixtral-8x22B, a MoE-based LLM with more than 140B parameters, the training time is limited to 8 hours. Detailed analysis shows that harder tasks, middle layers, and content words tend to activate more experts, providing valuable insights for future adaptive MoE system designs. Both the training code and model checkpoints will be publicly available.

📄 PDF Abstract BibTeX arXiv:2410.10456

Code (0)

등록된 구현이 없습니다.

Tasks

Computational EfficiencyMixture-of-Experts

Methods 이 논문이 사용한 방법론

MoE 설명 없음

Similar Papers 제목 키워드 기반

LoRA-Switch: Boosting the Efficiency of Dynamic LLM Adapters via System-Algorithm Co-design

2024-05-28 · Rui Kong, Qiyang Li, Xinyu Fang, Qingtian Feng 외

Recent literature has found that an effective method to customize or further improve large language models (LLMs) is to add dynamic adapters, such as low-rank adapters (LoRA) with Mixture-of-Experts (MoE) structures. Tho…

Mixture-of-Experts

A Novel Mobility Model to Support the Routing of Mobile Energy Resources

2020-07-22 · Wei Wang, Xiaofu Xiong, Chao Xiao, Bihui Wei

Mobile energy resources (MERs) have received increasing attention due to their effectiveness in boosting the power system resilience in a flexible way. In this paper, a novel mobility model for MERs is proposed, which ca…

Computational Efficiency

DynMoLE: Boosting Mixture of LoRA Experts Fine-Tuning with a Hybrid Routing Mechanism

2025-04-01 · Dengchun Li, Naizheng Wang, Zihao Zhang, Haoyang Yin 외

Instruction-based fine-tuning of large language models (LLMs) has achieved remarkable success in various natural language processing (NLP) tasks. Parameter-efficient fine-tuning (PEFT) methods, such as Mixture of LoRA Ex…

Common Sense ReasoningComputational EfficiencyMixture-of-Expertsparameter-efficient fine-tuning

MoEEdit: Efficient and Routing-Stable Knowledge Editing for Mixture-of-Experts LLMs

2026-02-11 · Yupu Gu, Rongzhe Wei, Andy Zhu, Pan Li arxiv

Knowledge editing (KE) enables precise modifications to factual content in large language models (LLMs). Existing KE methods are largely designed for dense architectures, limiting their applicability to the increasingly …

knowledge editing

Towards Generalized Routing: Model and Agent Orchestration for Adaptive and Efficient Inference

2025-09-09 · Xiyu Guo, Shan Wang, Chunfang Ji, Xuefeng Zhao 외 arxiv

The rapid advancement of large language models (LLMs) and domain-specific AI agents has greatly expanded the ecosystem of AI-powered services. User queries, however, are highly diverse and often span multiple domains and…

Intent Recognition