paper-with-me

홈 › Papers

MoE-Inference-Bench: Performance Evaluation of Mixture of Expert Large Language and Vision Models

2025-08-24 · Krishna Teja Chitty-Venkata, Sylvia Howland, Golara Azar, Daria Soboleva, Natalia Vassilieva, Siddhisanket Raskar, Murali Emani, Venkatram Vishwanath arxiv

Mixture of Experts (MoE) models have enabled the scaling of Large Language Models (LLMs) and Vision Language Models (VLMs) by achieving massive parameter counts while maintaining computational efficiency. However, MoEs introduce several inference-time challenges, including load imbalance across experts and the additional routing computational overhead. To address these challenges and fully harness the benefits of MoE, a systematic evaluation of hardware acceleration techniques is essential. We present MoE-Inference-Bench, a comprehensive study to evaluate MoE performance across diverse scenarios. We analyze the impact of batch size, sequence length, and critical MoE hyperparameters such as FFN dimensions and number of experts on throughput. We evaluate several optimization techniques on Nvidia H100 GPUs, including pruning, Fused MoE operations, speculative decoding, quantization, and various parallelization strategies. Our evaluation includes MoEs from the Mixtral, DeepSeek, OLMoE and Qwen families. The results reveal performance differences across configurations and provide insights for the efficient deployment of MoEs.

📄 PDF Abstract BibTeX arXiv:2508.17467

Code (0)

등록된 구현이 없습니다.

Tasks

Computational Efficiency

Similar Papers 제목 키워드 기반

MoQE: Improve Quantization Model performance via Mixture of Quantization Experts

2025-08-09 · Jinhao Zhang, Yunquan Zhang, Boyang Zhang, Zeyu Liu 외 arxiv

Quantization method plays a crucial role in improving model efficiency and reducing deployment costs, enabling the widespread application of deep learning models on resource-constrained devices. However, the quantization…

CoSMoEs: Compact Sparse Mixture of Experts

2025-02-28 · Patrick Huber, Akshat Shrivastava, Ernie Chang, Chinnadhurai Sankar 외

Sparse Mixture of Expert (MoE) models are popular foundational architectures at large scale, however, under-explored at smaller sizes. Here, we show how to enable Compact Sparse Mixture of Experts (CoSMoEs) for on-device…

Mixture-of-Experts

Elastic MoE: Unlocking the Inference-Time Scalability of Mixture-of-Experts

2025-09-26 · Naibin Gu, Zhenyu Zhang, Yuchen Feng, Yilong Chen 외 arxiv

Mixture-of-Experts (MoE) models typically fix the number of activated experts $k$ at both training and inference. However, real-world deployments often face heterogeneous hardware, fluctuating workloads, and diverse qual…

M2R2: Mixture of Multi-Rate Residuals for Efficient Transformer Inference

2025-02-04 · Nikhil Bhendawade, Mahyar Najibi, Devang Naik, Irina Belousova

Residual transformations enhance the representational depth and expressive power of large language models (LLMs). However, applying static residual transformations across all tokens in auto-regressive generation leads to…

Mixture-of-Experts

Doubly Sparse: Sparse Mixture of Sparse Experts for Efficient Softmax Inference

2019-01-30 · ICLR 2019 5 · Shun Liao, Ting Chen, Tian Lin, Denny Zhou 외

Computations for the softmax function are significantly expensive when the number of output classes is large. In this paper, we present a novel softmax inference speedup method, Doubly Sparse Softmax (DS-Softmax), that l…

image-classificationImage ClassificationLanguage ModelingLanguage Modelling+2