paper-with-me

홈 › Papers

Efficient Mixture-of-Experts LLM Inference with Apple Silicon NPUs

2026-04-20 · Afsara Benazir, Felix Xiaozhu Lin arxiv

Apple Neural Engine (ANE) is a dedicated neural processing unit (NPU) present in every Apple Silicon chip. Mixture-of-Experts (MoE) LLMs improve inference efficiency via sparse activation but are challenging for NPUs in three ways: expert routing is unpredictable and introduces dynamic tensor shapes that conflict with the shape-specific constraints of NPUs; several irregular operators, e.g., top-k, scatter/gather, etc., are not NPU-friendly; and launching many small expert kernels incurs substantial dispatch and synchronization overhead. NPUs are designed to offload AI compute from CPU and GPU; our goal is to enable such offloading for MoE inference, particularly during prefill, where long-context workloads consume substantial system resources. This paper presents NPUMoE, a runtime inference engine that accelerates MoE execution on Apple Silicon by offloading dense, static computation to NPU, while preserving a CPU/GPU fallback path for dynamic operations. NPUMoE uses offline calibration to estimate expert capacity and popularity that drives three key techniques: (1) Static tiers for expert capacity to address dynamic expert routing (2) Grouped expert execution to mitigate NPU concurrency limits (3) Load-aware expert compute graph residency to reduce CPU-NPU synchronization overhead. Experiments on Apple M-series devices using three representative MoE LLMs and four long-context workloads show that NPUMoE consistently outperforms baselines, reducing latency by 1.32x-5.55x, improving energy efficiency by 1.81x-7.37x, and reducing CPU-cycle usage by 1.78x-5.54x through effective NPU offloading.

📄 PDF Abstract BibTeX arXiv:2604.18788

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

BaseRT: Best-in-Class LLM Inference on Apple Silicon via Native Metal

2026-07-01 · Prabod Rathnayaka, Fabian Waschkowski, Lukas Wesemann arxiv

We present BaseRT, a native Metal inference runtime for large language models (LLMs) on Apple Silicon, and report the highest inference throughput on this hardware to date. Existing runtimes, including llama.cpp and MLX-…

Apple Intelligence Foundation Language Models: Tech Report 2025

2025-07-17 · Ethan Li, Anders Boesen Lindbo Larsen, Chen Zhang, Xiyou Zhou 외 arxiv

We introduce two multilingual, multimodal foundation language models that power Apple Intelligence features across Apple devices and services: i a 3B-parameter on-device model optimized for Apple silicon through architec…

Reinforcement Learning

Native LLM and MLLM Inference at Scale on Apple Silicon

2026-01-27 · Wayner Barrios arxiv

The growing adoption of Apple Silicon for machine learning development has created demand for efficient inference solutions that leverage its unique unified memory architecture. However, existing tools either lack native…

Systematic Optimization of Real-Time Diffusion Model Inference on Apple M3 Ultra

2026-02-10 · Yoichi Ochiai arxiv

While real-time image generation using diffusion models has advanced rapidly on NVIDIA GPUs, systematic optimization research on non-CUDA platforms such as Apple Silicon remains extremely limited. In this study, we condu…

Knowledge DistillationImage Generation

Benchmarking On-Device Machine Learning on Apple Silicon with MLX

2025-10-21 · Oluwaseun A. Ajayi, Ogundepo Odunayo arxiv

The recent widespread adoption of Large Language Models (LLMs) and machine learning in general has sparked research interest in exploring the possibilities of deploying these models on smaller devices such as laptops and…