paper-with-me

홈 › Papers

Mixture of Diverse Size Experts

2024-09-18 · Manxi Sun, Wei Liu, Jian Luan, Pengzhi Gao, Bin Wang

The Sparsely-Activated Mixture-of-Experts (MoE) has gained increasing popularity for scaling up large language models (LLMs) without exploding computational costs. Despite its success, the current design faces a challenge where all experts have the same size, limiting the ability of tokens to choose the experts with the most appropriate size for generating the next token. In this paper, we propose the Mixture of Diverse Size Experts (MoDSE), a new MoE architecture with layers designed to have experts of different sizes. Our analysis of difficult token generation tasks shows that experts of various sizes achieve better predictions, and the routing path of the experts tends to be stable after a training period. However, having experts of diverse sizes can lead to uneven workload distribution. To tackle this limitation, we introduce an expert-pair allocation strategy to evenly distribute the workload across multiple GPUs. Comprehensive evaluations across multiple benchmarks demonstrate the effectiveness of MoDSE, as it outperforms existing MoEs by allocating the parameter budget to experts adaptively while maintaining the same total parameter size and the number of experts.

📄 PDF Abstract BibTeX arXiv:2409.12210

Code (0)

등록된 구현이 없습니다.

Tasks

Mixture-of-Experts

Methods 이 논문이 사용한 방법론

MoE 설명 없음

Similar Papers 제목 키워드 기반

Graph Mixture of Experts: Learning on Large-Scale Graphs with Explicit Diversity Modeling

2023-09-21 · NeurIPS 2023 11

Graph neural networks (GNNs) have found extensive applications in learning from graph data. However, real-world graphs often possess diverse structures and comprise nodes and edges of varying types. To bolster the genera…

HMoE: Heterogeneous Mixture of Experts for Language Modeling

2024-08-20 · An Wang, Xingwu Sun, Ruobing Xie, Shuaipeng Li 외

Mixture of Experts (MoE) offers remarkable performance and computational efficiency by selectively activating subsets of model parameters. Traditionally, MoE models use homogeneous experts, each with identical capacity. …

Computational EfficiencyLanguage ModelingLanguage ModellingMixture-of-Experts

FlexMoRE: A Flexible Mixture of Rank-heterogeneous Experts for Efficient Federatedly-trained Large Language Models

2026-02-09 · Annemette Brok Pirchert, Jacob Nielsen, Mogens Henrik From, Lukas Galke Poech 외 arxiv

Recent advances in mixture-of-experts architectures have shown that individual experts models can be trained federatedly, i.e., in isolation from other experts by using a common base model to facilitate coordination. How…

Style Mixture of Experts for Expressive Text-To-Speech Synthesis

2024-06-05 · Ahad Jawaid, Shreeram Suresh Chandra, Junchen Lu, Berrak Sisman

Recent advances in style transfer text-to-speech (TTS) have improved the expressiveness of synthesized speech. However, encoding stylistic information (e.g., timbre, emotion, and prosody) from diverse and unseen referenc…

Mixture-of-ExpertsSpeech SynthesisStyle Transfertext-to-speech+2

SEKE: Specialised Experts for Keyword Extraction

2024-12-18 · Matej Martinc, Hanh Thi Hong Tran, Senja Pollak, Boshko Koloski

Keyword extraction involves identifying the most descriptive words in a document, allowing automatic categorisation and summarisation of large quantities of diverse textual data. Relying on the insight that real-world ke…

DescriptiveKeyword ExtractionMixture-of-Experts