paper-with-me

홈 › Papers

From Score Distributions to Balance: Plug-and-Play Mixture-of-Experts Routing

2025-09-29 · Rana Shahout, Colin Cai, Yilun Du, Minlan Yu, Michael Mitzenmacher arxiv

Mixture-of-Experts (MoE) models can scale parameter capacity by routing each token to a subset of experts through a learned gate function. While conditional routing reduces training costs, it shifts the burden on inference memory: expert parameters and activations consume memory, limiting the number of experts per device. As tokens are routed, some experts become overloaded while others are underutilized. Because experts are mapped to GPUs, this imbalance translates directly into degraded system performance in terms of latency, throughput, and cost. We present LASER, a plug-and-play, inference-time routing algorithm that balances load while preserving accuracy. LASER adapts to the shape of the gate's score distribution. When scores provide a clear preference, it routes to the strongest experts; when scores are more uniform, it broadens the set of viable experts and routes to the least-loaded among them. Because LASER relies only on gate scores from a trained model, it integrates directly into existing MoE inference pipelines without retraining or finetuning. We evaluate LASER on Mixtral-8x7B and DeepSeek-MoE-16b-chat across four datasets (ARC-Easy, ARC-Challenge, MMLU, and GSM8K). LASER improves load balancing, translating into lower latency and higher throughput, while keeping the accuracy changes negligible.

📄 PDF Abstract BibTeX arXiv:2510.03293

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Plug-And-Play Learned Gaussian-mixture Approximate Message Passing

2020-11-18 · Osman Musa, Peter Jung, Giuseppe Caire

Deep unfolding showed to be a very successful approach for accelerating and tuning classical signal processing algorithms. In this paper, we propose learned Gaussian-mixture AMP (L-GM-AMP) - a plug-and-play compressed se…

compressed sensingDenoising

A Replicate-and-Quantize Strategy for Plug-and-Play Load Balancing of Sparse Mixture-of-Experts LLMs

2026-02-23 · Zijie Liu, Jie Peng, Jinhao Duan, Zirui Liu 외 arxiv

Sparse Mixture-of-Experts (SMoE) architectures are increasingly used to scale large language models efficiently, delivering strong accuracy under fixed compute budgets. However, SMoE models often suffer from severe load …

SkewRoute: Training-Free LLM Routing for Knowledge Graph Retrieval-Augmented Generation via Score Skewness of Retrieved Context

2025-05-28 · Hairu Wang, Yuan Feng, Yukun Cao, Xike Xie 외

Large language models excel at many tasks but often incur high inference costs during deployment. To mitigate hallucination, many systems use a knowledge graph to enhance retrieval-augmented generation (KG-RAG). However,…

HallucinationRAGRetrievalRetrieval-augmented Generation

SAIP: A Plug-and-Play Scale-adaptive Module in Diffusion-based Inverse Problems

2025-09-29 · Lingyu Wang, Xiangming Meng arxiv

Solving inverse problems with diffusion models has shown promise in tasks such as image restoration. A common approach is to formulate the problem in a Bayesian framework and sample from the posterior by combining the pr…

Image Restoration

Normalized Wasserstein for Mixture Distributions With Applications in Adversarial Learning and Domain Adaptation

2019-10-01 · ICCV 2019 10 · Yogesh Balaji, Rama Chellappa, Soheil Feizi

Understanding proper distance measures between distributions is at the core of several learning tasks such as generative models, domain adaptation, clustering, etc. In this work, we focus on mixture distributions that ar…

ClusteringDomain Adaptation