paper-with-me

홈 › Papers

Three Phases of Expert Routing: How Load Balance Evolves During Mixture-of-Experts Training

2026-04-05 · Charafeddine Mouzouni arxiv

We model Mixture-of-Experts (MoE) token routing as a congestion game with a single effective parameter, the congestion coefficient gamma_eff, that quantifies the balance-quality tradeoff. Tracking gamma_eff across training checkpoints of two open-source MoE models, OLMoE-1B-7B (20 checkpoints, with dense sampling in the surge region) and OpenMoE-8B (6 checkpoints), reveals a three-phase trajectory: a surge phase where the router learns to balance load (gamma_eff: 14 to 36-39, peaking in the step 30K-40K region), a stabilization phase where experts specialize under steady balance (B_0: 2.4 to 2.3, steps 100K-400K), and a relaxation phase where the router trades balance for quality as experts differentiate (gamma_eff: 27 to 9, steps 400K-1.2M). This non-monotone trajectory, invisible to post-hoc analysis of converged models, reveals that early MoE training prioritizes balance while late training prioritizes quality. The theoretical framework is honest about its limits: the single-type equilibrium reduces to temperature-scaled softmax (held-out L1: MFG = 0.199 vs. softmax = 0.200). The game is not a better predictor; it reveals what the temperature means and, critically, how that temperature evolves. We complement the dynamics with an effective congestion decomposition, a multi-type extension that improves load prediction via token clustering on all 16 layers (mean: 30%), scope diagnostics (K/M, epsilon_l), and robustness verification across four independent quality estimators (r >= 0.89). All confidence intervals are from bootstrap resampling over 50 independent text batches.

📄 PDF Abstract BibTeX arXiv:2604.04230

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

FreeBalance: Pre-Routing Online Moe Load Balancing via Residual Workload Prediction

2026-08-14 · Pengfei Chen, Yize Wu, Shouxu Kuang, Ke Gao 외 arxiv

Load imbalance poses a major bottleneck to the efficiency of expert parallelism in distributed inference of Mixture-of-Experts (MoE) models. The most heavily loaded rank stalls global execution due to skewed routing dist…

A Replicate-and-Quantize Strategy for Plug-and-Play Load Balancing of Sparse Mixture-of-Experts LLMs

2026-02-23 · Zijie Liu, Jie Peng, Jinhao Duan, Zirui Liu 외 arxiv

Sparse Mixture-of-Experts (SMoE) architectures are increasingly used to scale large language models efficiently, delivering strong accuracy under fixed compute budgets. However, SMoE models often suffer from severe load …

ELDR: Expert-Locality-Aware Decode Routing for PD-Disaggregated MoE Serving

2026-07-01 · Sangjin Choi, Sukmin Cho, Yifan Xiong, Ziyue Yang 외 hf

In prefill-decode (PD) disaggregated LLM serving, each request is assigned to a decode worker after prefill. Existing decode routers balance only load; for mixture-of-experts (MoE) models this is incomplete: equally load…

Least-Loaded Expert Parallelism: Load Balancing An Imbalanced Mixture-of-Experts

2026-01-23 · Xuan-Phi Nguyen, Shrey Pandit, Austin Xu, Caiming Xiong 외 arxiv

Mixture-of-Experts (MoE) models are typically pre-trained with explicit load-balancing constraints to ensure statistically balanced expert routing. Despite this, we observe that even well-trained MoE models exhibit signi…

Auxiliary-Loss-Free Load Balancing Strategy for Mixture-of-Experts

2024-08-28 · Lean Wang, Huazuo Gao, Chenggang Zhao, Xu sun 외

For Mixture-of-Experts (MoE) models, an unbalanced expert load will lead to routing collapse or increased computational overhead. Existing methods commonly employ an auxiliary loss to encourage load balance, but a large …

Mixture-of-Experts