paper-with-me

Papers

Pipeline Parallelism is All You Need for Optimized Early-Exit Based Self-Speculative Decoding

2025-09-19 · Ruanjun Li, Ziheng Liu, Yuanming Shi, Jiawei Shao, Chi Zhang, Xuelong Li arxiv

Large language models (LLMs) deliver impressive generation quality, but incur very high inference cost because each output token is generated auto-regressively through all model layers. Early-exit based self-speculative decoding (EESD) has emerged to mitigate this cost. However, in practice, many approaches struggle to achieve the expected acceleration in such draft-then-verify paradigm even with a well-aligned early-exit head and selected exit position. Our analysis reveals that EESD only pays off when the vast majority of draft tokens are accepted by the LLM. Otherwise, the draft cost may overcome the acceleration gain and lead to a negative speedup. To mitigate this, we propose Pipeline-Parallel Self-Speculative Decoding (PPSD) that fully pipelines the draft and verification work so that no effort is wasted on failed predictions. It has two key innovations. We configure the model layers as a pipeline in which early-exit (draft) computations and remaining-layer (verification) computations overlap. We interleave drafting and verification per token. While the LLM is verifying the current token in its final layers, the early-exit path simultaneously drafts the next token. Such a verify-while-draft scheme keeps all units busy and validates tokens on-the-fly analogous to pipelining the speculation and verification stages. Empirical results confirm that PPSD achieves state-of-the-art acceleration in self-speculative LLM inference. On diverse benchmarks, PPSD achieves speedup ratios in the range of 2.01x~3.81x, which gains almost the optimal acceleration at the fixed acceptance rate and exit position, showcasing its advancement in providing efficient self-speculation.

📄 PDF Abstract BibTeX arXiv:2509.19368

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

EE-LLM: Large-Scale Training and Inference of Early-Exit Large Language Models with 3D Parallelism

2023-12-08 · Yanxi Chen, Xuchen Pan, Yaliang Li, Bolin Ding 외

We present EE-LLM, a framework for large-scale training and inference of early-exit large language models (LLMs). While recent works have shown preliminary evidence for the efficacy of early exiting in accelerating LLM i…

BaPipe: Exploration of Balanced Pipeline Parallelism for DNN Training

2020-12-23 · Letian Zhao, Rui Xu, Tianqi Wang, Teng Tian 외

The size of deep neural networks (DNNs) grows rapidly as the complexity of the machine learning algorithm increases. To satisfy the requirement of computation and memory of DNN training, distributed deep learning based o…

GPU

Pipeline Gradient-based Model Training on Analog In-memory Accelerators

2024-10-19 · Zhaoxian Wu, Quan Xiao, Tayfun Gokmen, Hsinyu Tsai 외

Aiming to accelerate the training of large deep neural models (DNN) in an energy-efficient way, an analog in-memory computing (AIMC) accelerator emerges as a solution with immense potential. In AIMC accelerators, trainab…

DeepSpeed Ulysses: System Optimizations for Enabling Training of Extreme Long Sequence Transformer Models

2023-09-25 · Sam Ade Jacobs, Masahiro Tanaka, Chengming Zhang, Minjia Zhang 외

Computation in a typical Transformer-based large language model (LLM) can be characterized by batch size, hidden dimension, number of layers, and sequence length. Until now, system works for accelerating LLM training hav…

Language ModellingLarge Language Model

Dynamic Decision Tree Ensembles for Energy-Efficient Inference on IoT Edge Nodes

2023-06-16 · Francesco Daghero, Alessio Burrello, Enrico Macii, Paolo Montuschi 외

With the increasing popularity of Internet of Things (IoT) devices, there is a growing need for energy-efficient Machine Learning (ML) models that can run on constrained edge nodes. Decision tree ensembles, such as Rando…

C++ code