paper-with-me

Papers

BudgetDraft: Acceptance-Aware Multi-View Training for Sparse-KV Speculative Decoding

2026-05-29 · Liang He, Jingbo Wen, Qishi Zhan, Yixiong Chen, Kangning Cui, Qizhen Lan, Xilu Wang arxiv

Speculative decoding speeds up autoregressive decoding by using a drafter to propose multiple tokens that a verifier validates in parallel. In resource-constrained deployments, the drafter uses a sparse KV cache to limit peak GPU memory and end-to-end latency under a fixed KV budget, while the verifier keeps a full KV cache. Mid-to-long context inference (4K--16K context length) is common in real applications. However, naive sparse/full speculative decoding suffers from the sparse/full mismatch as context length grows, causing the acceptance rate to drop quickly. We propose BudgetDraft, a multi-view sparse training method for sparse drafting in mid-to-long inference. The drafter is exposed to multiple sampled KV budgets during training and learns to align each sparse view with one shared full-cache teacher target. BudgetDraft combines an acceptance-aware loss on a full-cache branch with a multi-view loss on a sparse-cache branch, producing a single budget-robust drafter that recovers acceptance across sparsity levels without extra inference-time components. Experimental results on PG-19, LongBench, and LWM show that BudgetDraft achieves up to 6.55x, 4.46x, and 2.10x end-to-end speedup vs AR at 4K, 8K, and 16K context lengths, while keeping the inference pipeline memory-friendly.

📄 PDF Abstract BibTeX arXiv:2606.00144

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

UniScale: Unified Scale-Aware 3D Reconstruction for Multi-View Understanding via Prior Injection for Robotic Perception

2026-02-26 · Mohammad Mahdavian, Gordon Tan, Binbin Xu, Yuan Ren 외 arxiv

We present UniScale, a unified, scale-aware multi-view 3D reconstruction framework for robotic applications that flexibly integrates geometric priors through a modular, semantically informed design. In vision-based robot…

Multi-View 3D Reconstruction

Identity-Aware Human-Object Interaction Motion Captioning

2026-08-21 · Yiming Wang, Yonghao Dang, Huilai Li, Jiawei Tu 외 arxiv

Existing human-object interaction (HOI) motion captioning methods typically describe what happens while referring to the subject using generic terms such as "a person" or "someone", without grounding the caption in subje…

Motion Captioning

Communication, Awareness and Acceptance of Digital Banking Amidst Cash Crunch in Southeast and South-South, Nigeria

2025-04-13 · Okechukwu Christopher Onuegbu, Bettina Oboakore Agbamu, Belinda Uju Anyakoha, Ogonna Wilson Anunike

Digital banking is among the technological innovations currently reverberating the cyber wave. this study seeks to assess communication, awareness and acceptance of it among the residents of south-east and south-south, n…

Survey

DataEvolver: Let Your Data Build and Improve Itself via Goal-Driven Loop Agents

2026-05-03 · Qisong Zhang, Wenzhuo Wu, Zhuangzhuang Jia, Yunhao Yang 외 arxiv

Constructing controllable visual data is a major bottleneck for image editing and multimodal understanding. Useful supervision is rarely produced by a single rendering pass; instead it emerges through iterative generatio…

Image Editing

Generative Artificial Intelligence Adoption Among Bangladeshi Journalists: Exploring Journalists' Awareness, Acceptance, Usage, and Organizational Stance on Generative AI

2025-11-14 · H. M. Murtuza, Md Oliullah arxiv

Newsrooms and journalists across the world are adopting Generative AI (GenAI). Drawing on in-depth interviews with 23 journalists, this study identifies Bangladeshi journalists' awareness, acceptance, usage patterns, and…