paper-with-me

Papers

GrowMTP: Can RL Grow Its Own Draft Head?

2026-09-15 · Minghua He, Lingzhe Zhang, Yuan Liu, Xiao Zhou, Aiwei Liu arxiv

Reinforcement learning (RL) post-training drives the frontier capabilities of large language models, with its wall-clock dominated by autoregressive rollout generation. Speculative decoding is an established remedy for this bottleneck, but existing draft heads must be pretrained or warmed up before RL, introducing substantial training cost outside the RL run to be accelerated. We observe that RL training itself provides both conditions required for online draft-head training: its rollout distribution is far narrower than that of pretraining, and its verification step continuously produces supervision signals aligned with this distribution. Building on these observations, we propose GrowMTP, which uses this supervision to train a draft head from scratch entirely within the RL loop, with all head updates detached from the policy backbone. On Qwen3-4B (no draft head), MiMo-7B-SFT (weak head), and Qwen3.5-4B-Base (strong head), GrowMTP achieves rollout speedups of 2.13x, 1.93x, and 1.36x, and end-to-end speedups of 1.60x, 1.41x, and 1.20x, respectively. GrowMTP therefore serves existing RL training frameworks as a modular component, particularly offering a from-scratch acceleration path for models without pretrained draft heads.

📄 PDF Abstract BibTeX arXiv:2609.16648

Code (0)

등록된 구현이 없습니다.

Tasks

Reinforcement Learning

Similar Papers 제목 키워드 기반

JetSpec: Breaking the Scaling Ceiling of Speculative Decoding with Parallel Tree Drafting

2026-06-16 · Lanxiang Hu, Zhaoxiang Feng, Yulun Wu, Haoran Yuan 외 arxiv

Speculative decoding (SD) accelerates autoregressive Large Language Models (LLMs) by drafting multiple tokens and verifying them in parallel, but it faces a scaling limitation: increasing the draft budget improves speed …

Attention Drift: What Autoregressive Speculative Decoding Models Learn

2026-05-11 · Doğaç Eldenk, Payal Mohapatra, Yigitcan Comlek, Kaan Oktay 외 arxiv

Speculative decoding accelerates LLM inference by drafting future tokens with a small model, but drafter models degrade sharply under template perturbation and long-context inputs. We identify a previously-unreported phe…

Windowed-MTP: Removing the Full-Context Draft-KV Tax at Million-Token Context

2026-07-23 · Alagappan Valliappan arxiv

Speculative decoding accelerates autoregressive generation by having a cheap draft propose tokens that a target verifies in parallel. Frontier models increasingly ship a built-in Multi-Token-Prediction (MTP/NEXTN) draft …

Balancing Coverage and Draft Latency in Vocabulary Trimming for Faster Speculative Decoding

2026-03-05 · Ofir Ben Shoham arxiv

Speculative decoding accelerates inference for Large Language Models by using a lightweight draft model to propose candidate tokens that are verified in parallel by a larger target model. Prior work shows that the draft …

KOALA: Enhancing Speculative Decoding for LLM via Multi-Layer Draft Heads with Adversarial Learning

2024-08-15 · Kaiqi Zhang, Jing Zhao, Rui Chen

Large Language Models (LLMs) exhibit high inference latency due to their autoregressive decoding nature. While the draft head in speculative decoding mitigates this issue, its full potential remains unexplored. In this p…