paper-with-me

Papers

Bridging Draft Policy Misalignment: Group Tree Optimization for Speculative Decoding

2025-09-26 · Shijing Hu, Jingyang Li, Zhihui Lu, Pan Zhou arxiv

Speculative decoding accelerates large language model (LLM) inference by letting a lightweight draft model propose multiple tokens that the target model verifies in parallel. Yet existing training objectives optimize only a single greedy draft path, while decoding follows a tree policy that re-ranks and verifies multiple branches. This draft policy misalignment limits achievable speedups. We introduce Group Tree Optimization (GTO), which aligns training with the decoding-time tree policy through two components: (i) Draft Tree Reward, a sampling-free objective equal to the expected acceptance length of the draft tree under the target model, directly measuring decoding performance; (ii) Group-based Draft Policy Training, a stable optimization scheme that contrasts trees from the current and a frozen reference draft model, forming debiased group-standardized advantages and applying a PPO-style surrogate along the longest accepted sequence for robust updates. We further prove that increasing our Draft Tree Reward provably improves acceptance length and speedup. Across dialogue (MT-Bench), code (HumanEval), and math (GSM8K), and multiple LLMs (e.g., LLaMA-3.1-8B, LLaMA-3.3-70B, Vicuna-1.3-13B, DeepSeek-R1-Distill-LLaMA-8B, Qwen3-8B), GTO increases acceptance length by (7.4%) and yields an additional (7.7%) speedup over prior state-of-the-art EAGLE-3. By bridging draft policy misalignment, GTO offers a practical, general solution for efficient LLM inference. Code and draft models are available at https://github.com/hsj576/GTO.

📄 PDF Abstract BibTeX arXiv:2509.22134

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

CORAL: Learning Consistent Representations across Multi-step Training with Lighter Speculative Drafter

2025-02-24 · Yepeng Weng, Dianwen Mei, Huishi Qiu, Xujie Chen 외

Speculative decoding is a powerful technique that accelerates Large Language Model (LLM) inference by leveraging a lightweight speculative draft model. However, existing designs suffers in performance due to misalignment…

Large Language Model

Model Trees for Identifying Exceptional Players in the NHL Draft

2018-02-23 · Oliver Schulte, Yejia Liu, Chao Li

Drafting strong players is crucial for the team success. We describe a new data-driven interpretable approach for assessing draft prospects in the National Hockey League. Successful previous approaches have built a predi…

Yggdrasil: Bridging Dynamic Speculation and Static Runtime for Latency-Optimal Tree-Based LLM Decoding

2025-12-29 · Yue Guan, Changming Yu, Shihan Fang, Weiming Hu 외 arxiv

Speculative decoding improves LLM inference by generating and verifying multiple tokens in parallel, but existing systems suffer from suboptimal performance due to a mismatch between dynamic speculation and static runtim…

iGRPO: Self-Feedback-Driven LLM Reasoning

2026-02-09 · Ali Hatamizadeh, Shrimai Prabhumoye, Igor Gitman, Ximing Lu 외 arxiv

Large Language Models (LLMs) have shown promise in solving complex mathematical problems, yet they still fall short of producing accurate and consistent solutions. Reinforcement Learning (RL) is a framework for aligning …

Reinforcement LearningMathematical Reasoning

SRT: Accelerating Reinforcement Learning via Speculative Rollout with Tree-Structured Cache

2026-01-14 · Chi-Chih Chang, Siqi Zhu, Zhichen Zeng, Haibin Lin 외 arxiv

We present Speculative Rollout with Tree-Structured Cache (SRT), a simple, model-free approach to accelerate on-policy reinforcement learning (RL) for language models without sacrificing distributional correctness. SRT e…

Reinforcement Learning