paper-with-me

홈 › Papers

When NPUs Are Not Always Faster: A Stage-Level Analysis of Mobile LLM Inference

2026-05-22 · Pu Li, Jiawen Qi, Qinyu Chen arxiv

Deploying large language models (LLMs) on mobile devices increasingly relies on heterogeneous execution, yet no prior study has systematically characterized NPU effectiveness at the operator and pipeline level. We present the first stage-aware, multi-level benchmarking study of mobile LLM inference on a CPU-NPU heterogeneous SoC. We introduce an OPMASK-based controlled pipeline decomposition methodology that isolates communication, quantization, and computation overheads within the NPU execution path. Our results reveal a counter-intuitive stage-level performance reversal: CPUs outperform NPUs in the compute-intensive Prefill stage (up to 1.6x), while NPUs provide only limited acceleration in the memory-bound Decode stage (1.05-1.2x). We further show that scheduling overhead and cross-backend fallback reduce the practical benefits of NPU offloading. For the energy trend, increasing NPU offloading leads to higher energy consumption (up to 51%). Based on these findings, we derive design guidelines for NPU architects targeting on-device LLM inference.

📄 PDF Abstract BibTeX arXiv:2605.27435

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

FastAttention: Extend FlashAttention2 to NPUs and Low-resource GPUs

2024-10-22 · Haoran Lin, Xianzhi Yu, Kang Zhao, Lu Hou 외

FlashAttention series has been widely applied in the inference of large language models (LLMs). However, FlashAttention series only supports the high-level GPU architectures, e.g., Ampere and Hopper. At present, FlashAtt…

CPUGPU

Edge Deployment of Small Language Models, a comprehensive comparison of CPU, GPU and NPU backends

2025-11-27 · Pablo Prieto, Pablo Abad arxiv

Edge computing processes data where it is generated, enabling faster decisions, lower bandwidth usage, and improved privacy. However, edge devices typically operate under strict constraints on processing power, memory, a…

MindVL: Towards Efficient and Effective Training of Multimodal Large Language Models on Ascend NPUs

2025-09-15 · Feilong Chen, Yijiang Liu, Yi Huang, Hao Wang 외 arxiv

We propose MindVL, a multimodal large language model (MLLMs) trained on Ascend NPUs. The training of state-of-the-art MLLMs is often confined to a limited set of hardware platforms and relies heavily on massive, undisclo…

Fast On-device LLM Inference with NPUs

2024-07-08 · Daliang Xu, Hao Zhang, Liming Yang, Ruiqi Liu 외

On-device inference for Large Language Models (LLMs), driven by increasing privacy concerns and advancements of mobile-sized models, has gained significant interest. However, even mobile-sized LLMs (e.g., Gemma-2B) encou…

CPUGPU

Pangu Ultra MoE: How to Train Your Big MoE on Ascend NPUs

2025-05-07 · Yehui Tang, Yichun Yin, Yaoyuan Wang, Hang Zhou 외

Sparse large language models (LLMs) with Mixture of Experts (MoE) and close to a trillion parameters are dominating the realm of most capable language models. However, the massive model scale poses significant challenges…

Mixture-of-Experts