paper-with-me

홈 › Papers

Goose: Anisotropic Speculation Trees for Training-Free Speculative Decoding

2026-04-02 · Tao Jin, Phuong Minh Nguyen, Naoya Inoue arxiv

Speculative decoding accelerates large language model inference by drafting multiple candidate tokens and verifying them in a single forward pass. Candidates are organized as a tree: deeper trees accept more tokens per step, but adding depth requires sacrificing breadth (fallback options) under a fixed verification budget. Existing training-free methods draft from a single token source and shape their trees without distinguishing candidate quality across origins. We observe that two common training-free token sources - n-gram matches copied from the input context, and statistical predictions from prior forward passes - differ dramatically in acceptance rate (~6x median gap, range 2-18x across five models and five benchmarks). We prove that when such a quality gap exists, the optimal tree is anisotropic (asymmetric): reliable tokens should form a deep chain while unreliable tokens spread as wide branches, breaking through the depth limit of balanced trees. We realize this structure in GOOSE, a training-free framework that builds an adaptive spine tree - a deep chain of high-acceptance context-matched tokens with wide branches of low-acceptance alternatives at each node. We prove that the number of tokens accepted per step is at least as large as that of either source used alone. On five LLMs (7B-33B) and five benchmarks, GOOSE achieves 1.9-4.3x lossless speedup, outperforming balanced-tree baselines by 12-33% under the same budget.

📄 PDF Abstract BibTeX arXiv:2604.02047

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

An Exploration of Left-Corner Transformations

2023-11-27 · Andreas Opedal, Eleftheria Tsipidi, Tiago Pimentel, Ryan Cotterell 외

The left-corner transformation (Rosenkrantz and Lewis, 1970) is used to remove left recursion from context-free grammars, which is an important step towards making the grammar parsable top-down with simple techniques. Th…

EntMTP: Accelerating LLM Inference with Entropy Guided Multi Token Prediction

2026-06-25 · Carrie Chen arxiv

Multi-token prediction has been shown to increase data density during training, improve downstream text-generation quality, and serves as the defacto approach for self-speculative decoding. Existing foundation and open s…

MoE-Spec: Expert Budgeting for Efficient Speculative Decoding

2026-02-17 · Bradley McDanel, Steven Li, Sruthikesh Surineni, Harshit Khaitan arxiv

Speculative decoding accelerates Large Language Model (LLM) inference by verifying multiple drafted tokens in parallel. However, for Mixture-of-Experts (MoE) models, this parallelism introduces a severe bottleneck: large…

GOOSE-M2F: Adapting Mask2Former for High-Fidelity, Long-Tailed Fine-Grained Semantic Segmentation in Unstructured Outdoor Terrain

2026-06-14 · Jyothiraditya Lingam, Nikhileswara Rao Sulake, Sai Manikanta Eswar Machara arxiv

We present GOOSE-M2F, a task-specific adaptation of Mask2Former for the GOOSE 2D Fine-Grained Semantic Segmentation (FGSS) Challenge at ICRA 2026. The GOOSE benchmark spans 64 fine-grained classes across unstructured out…

Semantic Segmentation

Golden Goose: A Simple Trick to Synthesize Unlimited RLVR Tasks from Unverifiable Internet Text

2026-01-30 · Ximing Lu, David Acuna, Jaehun Jung, Jian Hu 외 arxiv

Reinforcement Learning with Verifiable Rewards (RLVR) has become a cornerstone for unlocking complex reasoning in Large Language Models (LLMs). Yet, scaling up RL is bottlenecked by limited existing verifiable data, wher…

Reinforcement Learning