paper-with-me

Papers

TABED: Test-Time Adaptive Ensemble Drafting for Robust Speculative Decoding in LVLMs

2026-01-28 · Minjae Lee, Wonjun Kang, Byeongkeun Ahn, Christian Classen, Kevin Galim, Seunghyuk Oh, Minghao Yan, Hyung Il Koo, Kangwook Lee arxiv

Speculative decoding (SD) has proven effective for accelerating LLM inference by quickly generating draft tokens and verifying them in parallel. However, SD remains largely unexplored for Large Vision-Language Models (LVLMs), which extend LLMs to process both image and text prompts. To address this gap, we benchmark existing inference methods with small draft models on 11 datasets across diverse input scenarios and observe scenario-specific performance fluctuations. Motivated by these findings, we propose Test-time Adaptive Batched Ensemble Drafting (TABED), which dynamically ensembles multiple drafts obtained via batch inference by leveraging deviations from past ground truths available in the SD setting. The dynamic ensemble method achieves an average robust walltime speedup of 1.74x over autoregressive decoding and a 5% improvement over single drafting methods, while remaining training-free and keeping ensembling costs negligible through parameter sharing. With its plug-and-play compatibility, we further enhance TABED by integrating advanced verification and alternative drafting methods. Code and custom-trained models are available at https://github.com/furiosa-ai/TABED.

📄 PDF Abstract BibTeX arXiv:2601.20357

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

AHASD: Asynchronous Heterogeneous Architecture for LLM Adaptive Drafting Speculative Decoding on Mobile Devices

2026-04-28 · Ma Zirui, Fan Zhihua, Li Wenxing, Wu Haibin 외 arxiv

Speculative decoding enhances the inference efficiency of large language models (LLMs) by generating drafts using a small draft language model (DLM) and verifying them in batches with a large target language model (TLM).…

Learning to Draft: Adaptive Speculative Decoding with Reinforcement Learning

2026-03-02 · Jiebin Zhang, Zhenghan Yu, Liang Wang, Nan Yang 외 arxiv

Speculative decoding accelerates large language model (LLM) inference by using a small draft model to generate candidate tokens for a larger target model to verify. The efficacy of this technique hinges on the trade-off …

Reinforcement Learning

SJD-PAC: Accelerating Speculative Jacobi Decoding via Proactive Drafting and Adaptive Continuation

2026-03-19 · Jialiang Kang, Han Shu, Wenshuo Li, Yingjie Zhai 외 arxiv

Speculative Jacobi Decoding (SJD) offers a draft-model-free approach to accelerate autoregressive text-to-image synthesis. However, the high-entropy nature of visual generation yields low draft-token acceptance rates in …

Accelerated Test-Time Scaling with Model-Free Speculative Sampling

2025-06-05 · Woomin Song, Saket Dingliwal, Sai Muralidhar Jayanthi, Bhavana Ganesh 외

Language models have demonstrated remarkable capabilities in reasoning tasks through test-time scaling techniques like best-of-N sampling and tree search. However, these approaches often demand substantial computational …

Language ModelingLanguage Modelling

Using Covid-19 Response Policy to Estimate Open Water Swim Drafting Effects in Triathlon

2025-02-13 · Felix Reichel

This study investigates the causal effects of open-water swim drafting by leveraging a natural experiment induced by staggered race starts during the COVID-19 pandemic. Before 2020, athletes started in groups, enabling d…