paper-with-me

홈 › Papers

Fast and Cost-effective Speculative Edge-Cloud Decoding with Early Exits

2025-05-27 · Yeshwanth Venkatesha, Souvik Kundu, Priyadarshini Panda

Large Language Models (LLMs) enable various applications on edge devices such as smartphones, wearables, and embodied robots. However, their deployment often depends on expensive cloud-based APIs, creating high operational costs, which limit access for smaller organizations and raise sustainability concerns. Certain LLMs can be deployed on-device, offering a cost-effective solution with reduced latency and improved privacy. Yet, limited computing resources constrain the size and accuracy of models that can be deployed, necessitating a collaborative design between edge and cloud. We propose a fast and cost-effective speculative edge-cloud decoding framework with a large target model on the server and a small draft model on the device. By introducing early exits in the target model, tokens are generated mid-verification, allowing the client to preemptively draft subsequent tokens before final verification, thus utilizing idle time and enhancing parallelism between edge and cloud. Using an NVIDIA Jetson Nano (client) and an A100 GPU (server) with Vicuna-68M (draft) and Llama2-7B (target) models, our method achieves up to a 35% reduction in latency compared to cloud-based autoregressive decoding, with an additional 11% improvement from preemptive drafting. To demonstrate real-world applicability, we deploy our method on the Unitree Go2 quadruped robot using Vision-Language Model (VLM) based control, achieving a 21% speedup over traditional cloud-based autoregressive decoding. These results demonstrate the potential of our framework for real-time LLM and VLM applications on resource-constrained edge devices.

📄 PDF Abstract BibTeX arXiv:2505.21594

Code (0)

등록된 구현이 없습니다.

Tasks

GPU

Similar Papers 제목 키워드 기반

ConfigSpec: Profiling-Based Configuration Selection for Distributed Edge--Cloud Speculative LLM Serving

2026-04-08 · Xiangchen Li, Saeid Ghafouri, Jiakun Fan, Babar Ali 외 arxiv

Speculative decoding enables collaborative Large Language Model (LLM) inference across cloud and edge by separating lightweight token drafting from heavyweight verification. While prior systems show performance and cost …

SPADE: Speculative Decoding for Precise and Low Cost Distributed Edge Cloud Inference

2026-08-13 · Divya Jyoti Bajpai, Kishan Kumar Upadhyay, Manjesh Kumar Hanawal arxiv

Large Language Models (LLMs) have achieved remarkable success in natural language understanding and generation, but their deployment is constrained by high computational demands. Deploying smaller LLMs directly on the ed…

Natural Language Understanding

DSD: A Distributed Speculative Decoding Solution for Edge-Cloud Agile Large Model Serving

2025-11-26 · Fengze Yu, Leshu Li, Brad McDanel, Sai Qian Zhang arxiv

Large language model (LLM) inference often suffers from high decoding latency and limited scalability across heterogeneous edge-cloud environments. Existing speculative decoding (SD) techniques accelerate token generatio…

S3D: A Simple and Cost-Effective Self-Speculative Decoding Scheme for Low-Memory GPUs

2024-05-30 · Wei Zhong, Manasa Bharadwaj

Speculative decoding (SD) has attracted a significant amount of research attention due to the substantial speedup it can achieve for LLM inference. However, despite the high speedups they offer, speculative decoding meth…

GPUQuantization

Speculative Policy Orchestration: A Latency-Resilient Framework for Cloud-Robotic Manipulation

2026-03-19 · Chanh Nguyen, Shutong Jin, Florian T. Pokorny, Erik Elmroth arxiv

Cloud robotics enables robots to offload high-dimensional motion planning and reasoning to remote servers. However, for continuous manipulation tasks requiring high-frequency control, network latency and jitter can sever…

Motion Planning