paper-with-me

홈 › Papers

Flow-Controlled Scheduling for LLM Inference with Provable Stability Guarantees

2026-04-13 · Zhuolun Dong, Junyu Cao arxiv

Large language models (LLMs) have been widely adopted due to their great performance across a wide range of applications. ChatGPT and Gemini now serve hundreds of millions of active users and handle billions of user requests per day, which puts optimizing LLM inference into the spotlight. A key challenge in LLM inference is that decode lengths are unknown. The memory usage for each request grows with generated tokens, which may lead to overflow and cause system instability. To address this concern, we propose a simple flow-control framework that controls the rate at which prompts join the active set. We derive a necessary condition that any stable system must satisfy and establish sufficient conditions under which our algorithm provably achieves stability. Experiments show that, compared to commonly used strategies in practice, our approach achieves higher token and request throughput, lower average and tail latency, and more stable KV cache utilization.

📄 PDF Abstract BibTeX arXiv:2604.11001

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

DynAMO:Dynamic Asset Management Orchestration via Topological Multi-Agent Scheduling

2026-06-14 · Kanishk Kushwaha, Vikrant Vinod Bansode, Harsh Vardhan, Dhaval C. Patel arxiv

While LLM-powered agents offer end-to-end automation for industrial asset lifecycles, real-world Industry 4.0 deployment is hindered by latency, concurrency instability, and safety risks. We present DynAMO (Dynamic Asset…

SAGA: Workflow-Atomic Scheduling for AI Agent Inference on GPU Clusters

2026-05-01 · Dongxin Guo, Jikun Wu, Siu Ming Yiu arxiv

AI agents execute tens to hundreds of chained LLM calls per task, yet GPU schedulers treat each call as independent, discarding gigabytes of intermediate state between steps and inflating end-to-end latency by 3-8x. We a…

Iterative Refinement of Flow Policies in Probability Space for Online Reinforcement Learning

2025-10-17 · Mingyang Sun, Pengxiang Ding, Weinan Zhang, Donglin Wang arxiv

While behavior cloning with flow/diffusion policies excels at learning complex skills from demonstrations, it remains vulnerable to distributional shift, and standard RL methods struggle to fine-tune these models due to …

Reinforcement Learning

ZipMoE: Efficient On-Device MoE Serving via Lossless Compression and Cache-Affinity Scheduling

2026-01-29 · Yuchen Yang, Yaru Zhao, Pu Yang, Shaowei Wang 외 arxiv

While Mixture-of-Experts (MoE) architectures substantially bolster the expressive power of large-language models, their prohibitive memory footprint severely impedes the practical deployment on resource-constrained edge …

MeanCache: From Instantaneous to Average Velocity for Accelerating Flow Matching Inference

2026-01-27 · Huanlin Gao, Ping Chen, Fuyuan Shi, Ruijia Wu 외 arxiv

We present MeanCache, a training-free caching framework for efficient Flow Matching inference. Existing caching methods reduce redundant computation but typically rely on instantaneous velocity information (e.g., feature…