paper-with-me

홈 › Papers

Is Your NPU Ready for LLMs? Dissecting the Hidden Efficiency Bottlenecks in Mobile LLM Inference

2026-07-06 · Guanyu Cai, Ruiming Tian, Lang Yang, Zhouhong Ren, Jinliang Yuan, Lingkun Li, Jiliang Wang arxiv

Deploying Large Language Models (LLMs) on mobile devices enhances privacy and reduces latency, but is severely bottlenecked by hardware inefficiency. We present the first comprehensive, cross-layer measurement study of mobile LLM inference, uniquely spanning five mainstream frameworks (e.g., llama.cpp, GENIE) and three hardware backends (CPU, GPU, NPU). To enable this analysis, we develop PowerBench, a fine-grained profiling tool that provides the first backend-specific energy attribution, moving beyond traditional device-level measurements. Our study yields three critical insights: (1) Framework-induced performance gaps are substantially amplified on NPUs, reaching up to 10x using custom operators due to divergent offloading and quantization strategies. (2) We identify a distinct phase split where NPUs excel at compute-bound prefilling, while CPUs outperform all other backends in memory-bound decoding. This is driven by the NPU's preference for large, fixed-shape workloads, which conflicts with the small-kernel, dynamic nature of decoding. (3) Backend-specific profiling uncovers substantial scheduling headroom missed by prior work. Suboptimal thread configurations, uncoordinated NPU sleep latencies, and CPU polling intervals result in up to 40% energy waste. Leveraging these findings, we present an energy-oriented best-practice configuration for mobile LLM inference. We estimate that this configuration could reduce energy consumption by up to 54.8% on the NPU backend across three datasets.

📄 PDF Abstract BibTeX arXiv:2607.05475

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Dissecting Transformers: A CLEAR Perspective towards Green AI

2025-10-03 · Hemang Jain, Shailender Goyal, Divyansh Pandey, Karthik Vaidhyanathan arxiv

The rapid adoption of Large Language Models (LLMs) has raised significant environmental concerns. Unlike the one-time cost of training, LLM inference occurs continuously and dominates the AI energy footprint. Yet most su…

The Hidden Cost of Readability: How Code Formatting Silently Consumes Your LLM Budget

2025-08-19 · Dangfeng Pan, Zhensu Sun, Cenyuan Zhang, David Lo 외 arxiv

Source code is usually formatted with elements like indentation and newlines to improve readability for human developers. However, these visual aids do not seem to be beneficial for large language models (LLMs) in the sa…

Code Completion

The LLM Already Knows: Estimating LLM-Perceived Question Difficulty via Hidden Representations

2025-09-16 · Yubo Zhu, Dongrui Liu, Zecheng Lin, Wei Tong 외 arxiv

Estimating the difficulty of input questions as perceived by large language models (LLMs) is essential for accurate performance evaluation and adaptive inference. Existing methods typically rely on repeated response samp…

Dissecting Multifractal detrended cross-correlation analysis

2024-06-09 · Borko Stosic, Tatijana Stosic

In this work we address the question of the Multifractal detrended cross-correlation analysis method that has been subject to some controversies since its inception almost two decades ago. To this end we propose several …

Time Series

KA2L: A Knowledge-Aware Active Learning Framework for LLMs

2026-03-18 · Haoxuan Yin, Bojian Liu, Chen Tang, Yangfan Wang 외 arxiv

Fine-tuning large language models (LLMs) with high-quality knowledge has been shown to enhance their performance effectively. However, there is a paucity of research on the depth of domain-specific knowledge comprehensio…

Active Learning