paper-with-me

홈 › Papers

PyroDash: Cost-Efficient Token-Level Small-Large Language Model Collaborative Inference

2026-07-22 · Niqi Lyu, Pengtao Shi, Wei Qiu, Jianlin Zhong, Sicong Xia, Jianyao Ma, Yicheng Ding arxiv

Large language models (LLMs) provide strong reasoning capabilities but are expensive to serve at scale, whereas small language models (SLMs) are cheaper but less reliable on difficult problems. We introduce PyroDash, a cost-aware framework for token-level SLM-LLM collaborative inference. During generation, the SLM decides whether to request assistance by emitting a control token. A Collaborate Engine then sends the query and partial reasoning trace to a frozen LLM for completion through a single handoff. The policy is internalized in the SLM, requiring neither a separate router, LLM retraining, nor access to LLM logits. PyroDash trains the SLM in three stages: control-token embedding learning, offloading-oriented supervised fine-tuning, and cost-aware alignment with Group Relative Policy Optimization. Its reward balances answer accuracy against inference cost normalized by LLM-only inference. Across five mathematical reasoning benchmarks, PyroDash supports different accuracy-cost operating points. With $λ=0.05$, it achieves 64.04 percent average accuracy, 6.36 percentage points above the LLM-only baseline, while reducing cost by 20.4 percent. With $λ=0.6$, it achieves 54.55 percent accuracy with a 1.90 percent LLM token ratio and 0.012 LLM calls per example, reducing total cost from USD 49.36 to USD 1.78. These results show that learned token-level handoffs can reduce LLM use while preserving strong reasoning performance.

📄 PDF Abstract BibTeX arXiv:2607.20327

Code (0)

등록된 구현이 없습니다.

Tasks

Mathematical Reasoning

Similar Papers 제목 키워드 기반

Sub-Token Routing for KV Cache Compression

2026-04-23 · Wei Jiang, Wei Wang arxiv

Transformer inference often requires a large KV cache, especially for long-context language modeling and multimodal generation. Existing compression methods usually reduce cache cost by selecting, evicting, quantizing, o…

multimodal generation

CITER: Collaborative Inference for Efficient Large Language Model Decoding with Token-Level Routing

2025-02-04 · Wenhao Zheng, Yixiao Chen, Weitong Zhang, Souvik Kundu 외

Large language models have achieved remarkable success in various tasks but suffer from high computational costs during inference, limiting their deployment in resource-constrained applications. To address this issue, we…

Collaborative InferenceLanguage ModelingLanguage ModellingLarge Language Model

Dual-Track CoT: Budget-Aware Stepwise Guidance for Small LMs

2026-04-27 · Sagnik Chatterjee, Atharva Patil, Sricharan Ramesh arxiv

Large Language Models (LLMs) solve many reasoning tasks via chain-of-thought (CoT) prompting, but smaller models (about 7 to 8B parameters) still struggle with multi-step reasoning under tight compute and token budgets. …

ReSET: Accurate Latency-Critical NVFP4 Reasoning via Step-Aware Temperature Scaling

2026-06-11 · Sihwa Lee, Janghwan Lee, Donghoon Yoo, Jae Gon Kim 외 arxiv

Large reasoning models (LRMs) improve complex problem-solving by generating long intermediate reasoning traces, but this substantially increases inference costs. NVFP4 inference offers a promising approach to reduce both…

Beyond Distillation: Task-level Mixture-of-Experts for Efficient Inference

2021-09-24 · Findings (EMNLP) 2021 11 · Sneha Kudugunta, Yanping Huang, Ankur Bapna, Maxim Krikun 외

Sparse Mixture-of-Experts (MoE) has been a successful approach for scaling multilingual translation models to billions of parameters without a proportional increase in training computation. However, MoE models are prohib…

Mixture-of-ExpertsSentence