paper-with-me

홈 › Papers

FastMoE: A Fast Mixture-of-Expert Training System

2021-03-24 · Jiaao He, Jiezhong Qiu, Aohan Zeng, Zhilin Yang, Jidong Zhai, Jie Tang

Mixture-of-Expert (MoE) presents a strong potential in enlarging the size of language model to trillions of parameters. However, training trillion-scale MoE requires algorithm and system co-design for a well-tuned high performance distributed training system. Unfortunately, the only existing platform that meets the requirements strongly depends on Google's hardware (TPU) and software (Mesh Tensorflow) stack, and is not open and available to the public, especially GPU and PyTorch communities. In this paper, we present FastMoE, a distributed MoE training system based on PyTorch with common accelerators. The system provides a hierarchical interface for both flexible model design and easy adaption to different applications, such as Transformer-XL and Megatron-LM. Different from direct implementation of MoE models using PyTorch, the training speed is highly optimized in FastMoE by sophisticated high-performance acceleration skills. The system supports placing different experts on multiple GPUs across multiple nodes, enabling enlarging the number of experts linearly against the number of GPUs. The source of FastMoE is available at https://github.com/laekov/fastmoe under Apache-2 license.

📄 PDF Abstract BibTeX arXiv:2103.13262

Code (3)

davidmrau/mixture-of-experts 공식 구현 pytorch
laekov/fastmoe 공식 구현 pytorch
Hanlard/M2M-by-fastmoe pytorch

Tasks

GPULanguage ModelingLanguage Modelling

Methods 이 논문이 사용한 방법론

102 Ways to Reach To Someone At Expedia by Phone: Step-by-Step Guide 설명 없음
Linear Layer A Linear Layer is a projection $\mathbf{XW + b}$.
Ways To Talk To Someone at Southwest Airlines Via Phone, Email, Or Chat Options: A Step by Step Guide 설명 없음
Softmax The Softmax output function transforms a previous layer's output into a vector of probabilities. It is commonly used for multiclass classification. Given an input vector $x$…
Multi-Head Attention 설명 없음
Attention 설명 없음
Adaptive Softmax Adaptive Softmax is a speedup technique for the computation of probability distributions over words. The adaptive softmax is…
Adaptive Input Representations Adaptive Input Embeddings extend the adaptive softmax to input word representations. The factorization assigns more…

Similar Papers 제목 키워드 기반

TA-MoE: Topology-Aware Large Scale Mixture-of-Expert Training

2023-02-20 · Chang Chen, Min Li, Zhihua Wu, dianhai yu 외

Sparsely gated Mixture-of-Expert (MoE) has demonstrated its effectiveness in scaling up deep neural networks to an extreme scale. Despite that numerous efforts have been made to improve the performance of MoE from the mo…

Lifelong Mixture of Variational Autoencoders

2021-07-09 · Fei Ye, Adrian G. Bors

In this paper, we propose an end-to-end lifelong learning mixture of experts. Each expert is implemented by a Variational Autoencoder (VAE). The experts in the mixture system are jointly trained by maximizing a mixture o…

Lifelong learningMixture-of-Experts

Variational Mixture of Gaussian Process Experts

2008-12-01 · NeurIPS 2008 12 · Chao Yuan, Claus Neubauer

Mixture of Gaussian processes models extended a single Gaussian process with ability of modeling multi-modal data and reduction of training complexity. Previous inference algorithms for these models are mostly based on G…

Gaussian ProcessesMixture-of-Experts

Fast Feedforward Networks

2023-08-28 · Peter Belcak, Roger Wattenhofer

We break the linear link between the layer size and its inference cost by introducing the fast feedforward (FFF) architecture, a log-time alternative to feedforward networks. We demonstrate that FFFs are up to 220x faste…

Mixture-of-Experts

Accelerating Mixture-of-Experts Training with Adaptive Expert Replication

2025-04-28 · Athinagoras Skiadopoulos, Mark Zhao, Swapnil Gandhi, Thomas Norrie 외

Mixture-of-Experts (MoE) models have become a widely adopted solution to continue scaling model sizes without a corresponding linear increase in compute. During MoE model training, each input token is dynamically routed …

GPUMixture-of-Experts