paper-with-me

Papers

OmniDraft: A Cross-vocabulary, Online Adaptive Drafter for On-device Speculative Decoding

2025-07-03 · Ramchalam Kinattinkara Ramakrishnan, Zhaocong Yuan, Shaojie Zhuo, Chen Feng, Yicheng Lin, Chenzheng Su, Xiaopeng Zhang arxiv

Speculative decoding generally dictates having a small, efficient draft model that is either pretrained or distilled offline to a particular target model series, for instance, Llama or Qwen models. However, within online deployment settings, there are two major challenges: 1) usage of a target model that is incompatible with the draft model; 2) expectation of latency improvements over usage and time. In this work, we propose OmniDraft, a unified framework that enables a single draft model to operate with any target model and adapt dynamically to user data. We introduce an online n-gram cache with hybrid distillation fine-tuning to address the cross-vocabulary mismatch across draft and target models; and further improve decoding speed by leveraging adaptive drafting techniques. OmniDraft is particularly suitable for on-device LLM applications where model cost, efficiency and user customization are the major points of contention. This further highlights the need to tackle the above challenges and motivates the \textit{``one drafter for all''} paradigm. We showcase the proficiency of the OmniDraft framework by performing online learning on math reasoning, coding and text generation tasks. Notably, OmniDraft enables a single Llama-68M model to pair with various target models including Vicuna-7B, Qwen2-7B and Llama3-8B models for speculative decoding; and additionally provides up to 1.5-2x speedup.

📄 PDF Abstract BibTeX arXiv:2507.02659

Code (0)

등록된 구현이 없습니다.

Tasks

Text Generation

Similar Papers 제목 키워드 기반

Out-of-Vocabulary Sampling Boosts Speculative Decoding

2025-06-02 · Nadav Timor, Jonathan Mamou, Oren Pereg, Hongyang Zhang 외

Speculative decoding relies on fast and accurate drafters. Recent state-of-the-art language models employ larger and larger vocabularies, which significantly slows down drafters. One promising approach to boost the effic…

SlimSpec: Low-Rank Draft LM-Head for Accelerated Speculative Decoding

2026-05-11 · Anton Plaksin, Sergei Krutikov, Sergei Skvortsov, Alexander Samarin arxiv

Speculative decoding speeds up autoregressive generation in Large Language Models (LLMs) through a two-step procedure, where a lightweight draft model proposes tokens which the target model then verifies in a single forw…

DynaSpec: Context-aware Dynamic Speculative Sampling for Large-Vocabulary Language Models

2025-10-11 · Jinbin Zhang, Nasib Ullah, Erik Schultheis, Rohit Babbar arxiv

Speculative decoding accelerates LLM inference by letting a small drafter propose multiple tokens which a large target model verifies once per speculation step. As vocabularies scale past 10e5 tokens,verification cost in…

Accelerating LLM Inference with Lossless Speculative Decoding Algorithms for Heterogeneous Vocabularies

2025-01-31 · Nadav Timor, Jonathan Mamou, Daniel Korat, Moshe Berchansky 외

Accelerating the inference of large language models (LLMs) is a critical challenge in generative AI. Speculative decoding (SD) methods offer substantial efficiency gains by generating multiple tokens using a single targe…

TreeGraft: Adaptive Multi-Drafter Grafting for Tree-Based Speculative Decoding

2026-05-28 · Jiaming Fan, Daming Cao, Canchen Huang, Jiale Fu 외 arxiv

Speculative decoding accelerates large language model inference through a draft-then-verify paradigm. Building on this, tree-structured methods improve inference by organizing proposals into multiple candidate paths, inc…