paper-with-me

홈 › Papers

Efficient LLM inference solution on Intel GPU

2023-12-19 · Hui Wu, Yi Gan, Feng Yuan, Jing Ma, Wei Zhu, Yutao Xu, Hong Zhu, Yuhua Zhu, Xiaoli Liu, Jinghui Gu, Peng Zhao

Transformer based Large Language Models (LLMs) have been widely used in many fields, and the efficiency of LLM inference becomes hot topic in real applications. However, LLMs are usually complicatedly designed in model structure with massive operations and perform inference in the auto-regressive mode, making it a challenging task to design a system with high efficiency. In this paper, we propose an efficient LLM inference solution with low latency and high throughput. Firstly, we simplify the LLM decoder layer by fusing data movement and element-wise operations to reduce the memory access frequency and lower system latency. We also propose a segment KV cache policy to keep key/value of the request and response tokens in separate physical memory for effective device memory management, helping enlarge the runtime batch size and improve system throughput. A customized Scaled-Dot-Product-Attention kernel is designed to match our fusion policy based on the segment KV cache solution. We implement our LLM inference solution on Intel GPU and publish it publicly. Compared with the standard HuggingFace implementation, the proposed solution achieves up to 7x lower token latency and 27x higher throughput for some popular LLMs on Intel GPU.

📄 PDF Abstract BibTeX arXiv:2401.05391

Code (0)

등록된 구현이 없습니다.

Tasks

DecoderGPUManagement

Similar Papers 제목 키워드 기반

On Achieving Privacy-Preserving State-of-the-Art Edge Intelligence

2023-02-10 · Daphnee Chabal, Dolly Sapra, Zoltán Ádám Mann

Deep Neural Network (DNN) Inference in Edge Computing, often called Edge Intelligence, requires solutions to insure that sensitive data confidentiality and intellectual property are not revealed in the process. Privacy-p…

Edge-computingModel CompressionPositionPrivacy Preserving

Ultimate Intelligence Part II: Physical Measure and Complexity of Intelligence

2015-04-09 · Eray Özkural

We continue our analysis of volume and energy measures that are appropriate for quantifying inductive inference systems. We extend logical depth and conceptual jump size measures in AIT to stochastic problems, and physic…

On Solving a Stochastic Shortest-Path Markov Decision Process as Probabilistic Inference

2021-09-13 · Mohamed Baioumy, Bruno Lacerda, Paul Duckworth, Nick Hawes

Previous work on planning as active inference addresses finite horizon problems and solutions valid for online planning. We propose solving the general Stochastic Shortest-Path Markov Decision Process (SSP MDP) as probab…

valid

Network Analysis for Explanation

2017-12-07 · Hiroshi Kuwajima, Masayuki Tanaka

Safety critical systems strongly require the quality aspects of artificial intelligence including explainability. In this paper, we analyzed a trained network to extract features which mainly contribute the inference. Ba…

Designing Ecosystems of Intelligence from First Principles

2022-12-02 · Karl J Friston, Maxwell J D Ramstead, Alex B Kiefer, Alexander Tschantz 외

This white paper lays out a vision of research and development in the field of artificial intelligence for the next decade (and beyond). Its denouement is a cyber-physical ecosystem of natural and synthetic sense-making,…

Model Selection