paper-with-me

Papers

Approximate Speculative Decoding

2026-08-04 · Yuannuo Feng, Zegang Peng, Yuxin Xie, Yubing Ye, Yizhe Chen, Wenshuai Yao, Wenyong Zhou, Wang Kang arxiv

Speculative decoding accelerates autoregressive generation by verifying a draft block with a target model in parallel. Under standard greedy verification, decoding stops at the first draft token that differs from the target argmax, discarding the remaining target-scored suffix. Although accepting such a mismatch changes the decoding trajectory, it can make a contiguous suffix reusable when its tokens remain target-greedy under the realized prefix. In this paper, we introduce \textbf{Approximate Speculative Decoding (ASD)}, a training-free verifier that replaces binary first-mismatch truncation with budgeted longest-prefix selection. ASD accepts selected mismatches subject to a local target-logit regret gate, a per-block exception cap, and a persistent request-level regret budget, then reuses the contiguous target-greedy suffix without additional approximate decisions or target-model forward passes. ASD requires neither a new draft model nor fine-tuning, and exactly reduces to standard greedy verification when the budget is zero. Experiments show that ASD improves fixed-workload throughput by $3.05\%$--$15.26\%$ over matched strict verification and averages a $7.78\%$ gain across seven Qwen3-14B + DSpark-14B tasks. On DeepSeek-V4-Flash (284B) with DSpark it also raises verifier-side acceptance by roughly $10\%$--$16\%$ on GSM8K and MATH-500 in an FP4-to-FP8 compatibility setting. The source code is publicly available at: https://github.com/Kissmetothemoon/ASD

📄 PDF Abstract BibTeX arXiv:2608.03447

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Accelerating PayPal's Commerce Agent with Speculative Decoding: An Empirical Study on EAGLE3 with Fine-Tuned Nemotron Models

2026-03-27 · Ally Qin, Jian Wan, Sarat Mudunuri, Srinivasan Manoharan arxiv

We evaluate speculative decoding with EAGLE3 as an inference-time optimization for PayPal's Commerce Agent, powered by a fine-tuned llama3.1-nemotron-nano-8B-v1 model. Building on prior work (NEMO-4-PAYPAL) that reduced …

Mistletoe: Stealthy Acceleration-Collapse Attacks on Speculative Decoding

2026-05-13 · Shuoyang Sun, Chang Dai, Hao Fang, Kuofeng Gao 외 arxiv

Speculative decoding has become a widely adopted technique for accelerating large language model (LLM) inference by drafting multiple candidate tokens and verifying them with a target model in parallel. Its efficiency, h…

Optimized Multi-Token Joint Decoding with Auxiliary Model for LLM Inference

2024-07-12 · Zongyue Qin, Ziniu Hu, Zifan He, Neha Prakriya 외

Large language models (LLMs) have achieved remarkable success across diverse tasks, yet their inference processes are hindered by substantial time and energy demands due to single-token generation at each decoding step. …

Language ModellingLarge Language Model

Fast Inference from Transformers via Speculative Decoding

2022-11-30 · Yaniv Leviathan, Matan Kalman, Yossi Matias

Inference from large autoregressive models like Transformers is slow - decoding K tokens takes K serial runs of the model. In this work we introduce speculative decoding - an algorithm to sample from autoregressive model…

Language ModelingLanguage Modelling

Speculative Decoding in Decentralized LLM Inference: Turning Communication Latency into Computation Throughput

2025-11-13 · Jingwei Song, Wanyi Chen, Xinyuan Song, Max 외 arxiv

Speculative decoding accelerates large language model (LLM) inference by using a lightweight draft model to propose tokens that are later verified by a stronger target model. While effective in centralized systems, its b…