paper-with-me

Papers

Pre-Attention Expert Prediction and Prefetching for Mixture-of-Experts Large Language Models

2025-11-10 · Shien Zhu, Samuel Bohl, Robin Oester, Gustavo Alonso arxiv

Mixture-of-Experts (MoE) Large Language Models (LLMs) efficiently scale-up the model while keeping relatively low inference cost. As MoE models only activate part of the experts, related work has proposed expert prediction and caching methods to prefetch the experts for faster inference. However, existing approaches utilize the activations from the previous layer for prediction, incurring low accuracy and leave the first layer unoptimized. Applying complex layers or even training standalone networks for better prediction introduces high computation overhead. In this paper, we propose pre-attention expert prediction to achieve accurate and lightweight expert prefetching. The key insight is that some functions in LLMs are ranking-preserving, indicating that matching the ranking of selected experts using simple linear functions is possible. Therefore, we utilize the activations before the attention block in the same layer with 2 linear functions and ranking-aware loss to achieve accurate prediction, which also supports prefetching in the first layer. Our lightweight, pre-attention expert routers achieve 93.03% accuracy on DeepSeek V2 Lite, 94.69% on Qwen3-30B, and 97.62% on Phi-mini-MoE, showing about 15% improvement on absolute accuracy over the state-of-the-art methods.

📄 PDF Abstract BibTeX arXiv:2511.10676

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

MoBiLE: Efficient Mixture-of-Experts Inference on Consumer GPU with Mixture of Big Little Experts

2025-10-14 · Yushu Zhao, Yubin Qin, Yang Wang, Xiaolong Yang 외 arxiv

Mixture-of-Experts (MoE) models have recently demonstrated exceptional performance across a diverse range of applications. The principle of sparse activation in MoE models facilitates an offloading strategy, wherein acti…

BuddyMoE: Exploiting Expert Redundancy to Accelerate Memory-Constrained Mixture-of-Experts Inference

2025-11-13 · Yun Wang, Lingyun Yang, Senhao Yu, Yixiao Wang 외 arxiv

Mixture-of-Experts (MoE) architectures scale language models by activating only a subset of specialized expert networks for each input token, thereby reducing the number of floating-point operations. However, the growing…

A Spatio-Temporal Expert Prefetching Framework for Efficient MoE-based LLM Inference

2026-06-13 · Yingnan Zhao, Razvan Bunescu, Ahmed Louri, Avinash Karanth 외 arxiv

Mixture-of-Experts (MoE) based large language models (LLMs), such as Qwen and DeepSeek, have recently emerged as an effective approach to improving model capacity without proportionally increasing computational cost. By …

Code Generation

WiSP: A Working-Set View of Mixture-of-Experts Serving on Extremely Low-Resource Hardware

2026-06-20 · Jiamu Zhang, Liang Wu, Mayank Darbari, Liangjie Hong arxiv

Modern Mixture-of-Experts (MoE) models place most of their parameters in expert layers, yet only a small fraction of those experts are used for any token. The unused weights must still be stored where the GPU can reach t…

Fate: Fast Edge Inference of Mixture-of-Experts Models via Cross-Layer Gate

2025-02-17 · Zhiyuan Fang, Zicong Hong, Yuegui Huang, Yufeng Lyu 외

Large Language Models (LLMs) have demonstrated impressive performance across various tasks, and their application in edge scenarios has attracted significant attention. However, sparse-activated Mixture-of-Experts (MoE) …

GPUMixture-of-ExpertsQuantization