paper-with-me

홈 › Papers

On the Representation Collapse of Sparse Mixture of Experts

2022-04-20 · Zewen Chi, Li Dong, Shaohan Huang, Damai Dai, Shuming Ma, Barun Patra, Saksham Singhal, Payal Bajaj, Xia Song, Xian-Ling Mao, Heyan Huang, Furu Wei

Sparse mixture of experts provides larger model capacity while requiring a constant computational overhead. It employs the routing mechanism to distribute input tokens to the best-matched experts according to their hidden representations. However, learning such a routing mechanism encourages token clustering around expert centroids, implying a trend toward representation collapse. In this work, we propose to estimate the routing scores between tokens and experts on a low-dimensional hypersphere. We conduct extensive experiments on cross-lingual language model pre-training and fine-tuning on downstream tasks. Experimental results across seven multilingual benchmarks show that our method achieves consistent gains. We also present a comprehensive analysis on the representation and routing behaviors of our models. Our method alleviates the representation collapse issue and achieves more consistent routing than the baseline mixture-of-experts methods.

📄 PDF Abstract BibTeX arXiv:2204.09179

Code (2)

microsoft/unilm 공식 구현 pytorch
microsoft/torchscale pytorch

Tasks

ClusteringLanguage ModelingLanguage ModellingMixture-of-Experts

Similar Papers 제목 키워드 기반

S2MoE: Robust Sparse Mixture of Experts via Stochastic Learning

2025-03-29 · Giang Do, Hung Le, Truyen Tran

Sparse Mixture of Experts (SMoE) enables efficient training of large language models by routing input tokens to a select number of experts. However, training SMoE remains challenging due to the issue of representation co…

Mixture-of-Experts

SimSMoE: Solving Representational Collapse via Similarity Measure

2024-06-22 · Giang Do, Hung Le, Truyen Tran

Sparse mixture of experts (SMoE) have emerged as an effective approach for scaling large language models while keeping a constant computational cost. Regardless of several notable successes of SMoE, effective training su…

Mixture-of-Experts

On the effectiveness of discrete representations in sparse mixture of experts

2024-11-28 · Giang Do, Kha Pham, Hung Le, Truyen Tran

Sparse mixture of experts (SMoE) is an effective solution for scaling up model capacity without increasing the computational costs. A crucial component of SMoE is the router, responsible for directing the input to releva…

Mixture-of-ExpertsQuantization

CompeteSMoE - Effective Training of Sparse Mixture of Experts via Competition

2024-02-04 · Quang Pham, Giang Do, Huy Nguyen, TrungTin Nguyen 외

Sparse mixture of experts (SMoE) offers an appealing solution to scale up the model complexity beyond the mean of increasing the network's depth or width. However, effective training of SMoE has proven to be challenging …

Mixture-of-Experts

Hierarchical Mixture-of-Experts with Two-Stage Optimization

2026-05-08 · Gleb Molodtsov, Alexander Miasnikov, Aleksandr Beznosikov arxiv

Sparse Mixture-of-Experts (MoE) models scale capacity by routing each token to a small subset of experts. However, their routers exhibit a fundamental trade-off: strong load balancing can suppress expert specialization, …