paper-with-me

홈 › Papers

Routing-Free Mixture-of-Experts

2026-04-01 · Yilun Liu, Jinru Han, Sikuan Yan, Volker Tresp, Yunpu Ma arxiv

Standard Mixture-of-Experts (MoE) models rely on centralized routing mechanisms that introduce rigid inductive biases. We propose Routing-Free MoE which eliminates any hard-coded centralized designs including external routers, Softmax, Top-K and load balancing, instead encapsulating all activation functionalities within individual experts and directly optimized through continuous gradient flow, enabling each expert to determine its activation entirely on its own. We introduce a unified adaptive load-balancing framework to simultaneously optimize both expert-balancing and token-balancing objectives through a configurable interpolation, allowing flexible and customizable resource allocation. Extensive experiments show that Routing-Free MoE can consistently outperform baselines with better scalability and robustness. We analyze its behavior in detail and offer insights that may facilitate future MoE design ad optimization.

📄 PDF Abstract BibTeX arXiv:2604.00801

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

StableMoE: Stable Routing Strategy for Mixture of Experts

2021-11-16 · ACL ARR November 2021 11 · Anonymous

The Mixture-of-Experts (MoE) technique can scale up the model size of Transformers with an affordable computational overhead. We point out that existing learning-to-route MoE methods suffer from the routing fluctuation i…

Language ModelingLanguage ModellingMachine TranslationMixture-of-Experts

StableMoE: Stable Routing Strategy for Mixture of Experts

2022-04-18 · ACL 2022 5 · Damai Dai, Li Dong, Shuming Ma, Bo Zheng 외

The Mixture-of-Experts (MoE) technique can scale up the model size of Transformers with an affordable computational overhead. We point out that existing learning-to-route MoE methods suffer from the routing fluctuation i…

Language ModelingLanguage ModellingMachine TranslationMixture-of-Experts

Cosine-Similarity Routing with Semantic Anchors for Interpretable Mixture-of-Experts Language Models

2025-09-12 · Ivan Ternovtsii, Yurii Bilak arxiv

Mixture-of-Experts (MoE) models improve efficiency through sparse activation, but their learned gating functions provide limited insight into routing decisions. This work introduces the Semantic Resonance Architecture (S…

Auxiliary-Loss-Free Load Balancing Strategy for Mixture-of-Experts

2024-08-28 · Lean Wang, Huazuo Gao, Chenggang Zhao, Xu sun 외

For Mixture-of-Experts (MoE) models, an unbalanced expert load will lead to routing collapse or increased computational overhead. Existing methods commonly employ an auxiliary loss to encourage load balance, but a large …

Mixture-of-Experts

HyperRouter: Towards Efficient Training and Inference of Sparse Mixture of Experts

2023-12-12 · Giang Do, Khiem Le, Quang Pham, TrungTin Nguyen 외

By routing input tokens to only a few split experts, Sparse Mixture-of-Experts has enabled efficient training of large language models. Recent findings suggest that fixing the routers can achieve competitive performance …

Mixture-of-Experts