paper-with-me

홈 › Papers

Routing Manifold Alignment Improves Generalization of Mixture-of-Experts LLMs

2025-11-10 · Zhongyang Li, Ziyue Li, Tianyi Zhou arxiv

Sparse Mixture-of-Experts (MoE) have been widely adopted in recent large language models since it can efficiently scale up the model capability without increasing the inference cost. However, evaluations on broad downstream tasks reveal a consistent suboptimality of the routers in existing MoE LLMs, which results in a severe performance gap (e.g., 10-20% in accuracy) to the optimal routing. In this paper, we show that aligning the manifold of routing weights with that of task embedding can effectively reduce the gap and improve MoE LLMs' generalization performance. Our method, "Routing Manifold Alignment (RoMA)", introduces an additional manifold regularization term in the post-training objective and only requires lightweight finetuning of routers (with other parameters frozen). Specifically, the regularization encourages the routing weights of each sample to be close to those of its successful neighbors (whose routing weights lead to correct answers) in a task embedding space. Consequently, samples targeting similar tasks will share similar expert choices across layers. Building such bindings between tasks and experts over different samples is essential to achieve better generalization. Moreover, RoMA demonstrates the advantage of unifying the task understanding (by embedding models) with solution generation (by MoE LLMs). In experiments, we finetune routers in OLMoE, DeepSeekMoE, and Qwen3-MoE using RoMA. Evaluations on diverse benchmarks and extensive comparisons with baselines show the substantial improvement brought by RoMA.

📄 PDF Abstract BibTeX arXiv:2511.07419

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

TrueMoE: Dual-Routing Mixture of Discriminative Experts for Synthetic Image Detection

2025-09-19 · Laixin Zhang, Shuaibo Li, Wei Ma, Hongbin Zha arxiv

The rapid progress of generative models has made synthetic image detection an increasingly critical task. Most existing approaches attempt to construct a single, universal discriminative space to separate real from fake …

Generalization and Scaling Laws for Mixture-of-Experts Transformers

2026-04-10 · Mansour Zoubeirou a Mayaki arxiv

We develop a theory of generalization and scaling for Mixture-of-Experts (MoE) Transformers that cleanly separates \emph{active} per-input capacity from routing combinatorics. By conditioning on fixed routing patterns an…

RASA: Routing-Aware Safety Alignment for Mixture-of-Experts Models

2026-02-04 · Jiacheng Liang, Yuhui Wang, Tanqiu Jiang, Ting Wang arxiv

Mixture-of-Experts (MoE) language models introduce unique challenges for safety alignment due to their sparse routing mechanisms, which can enable degenerate optimization behaviors under standard full-parameter fine-tuni…

Routing on the Stiefel Manifold: When Does Adaptive Subspace Selection Help for Cross-Domain EEG Decoding?

2026-05-29 · Isabella Costa Maia, Pedro L. C. Rodrigues, Salem Said, Marco Congedo arxiv

Cross-domain EEG decoding remains challenging despite advances in Riemannian deep learning: covariance matrices from different subjects occupy systematically distinct regions of the SPD manifold, yet existing domain adap…

Domain AdaptationEeg Decoding

Value-and-Structure Alignment for Routing-Consistent Quantization of Mixture-of-Experts Models

2026-06-04 · Hancheol Park, Geonho Lee, Tairen Piao, Tae-Ho Kim arxiv

Mixture-of-Experts (MoE) models scale foundation models efficiently by activating only a subset of experts for each token, but their large number of expert parameters still makes quantization essential for practical depl…