paper-with-me

Papers

SAGA: Workflow-Atomic Scheduling for AI Agent Inference on GPU Clusters

2026-05-01 · Dongxin Guo, Jikun Wu, Siu Ming Yiu arxiv

AI agents execute tens to hundreds of chained LLM calls per task, yet GPU schedulers treat each call as independent, discarding gigabytes of intermediate state between steps and inflating end-to-end latency by 3-8x. We argue that this request-level abstraction is fundamentally mismatched to compound AI workloads, and propose a shift to program-level scheduling: treating the entire agent workflow (not individual inference calls) as the first-class schedulable unit. We present SAGA, a distributed scheduler that implements this abstraction through three mechanisms: (1) Agent Execution Graphs that capture workflow structure to predict KV cache reuse across tool-call boundaries, achieving within 1.31x of Bélády's optimal offline policy; (2) session-affinity batching with work stealing that co-locates correlated requests while maintaining global load balance; and (3) Agent Fair Share, a task-completion-time fairness metric with provable bounded-deviation guarantees. On a 64-GPU cluster serving SWE-bench coding agents and WebArena browser tasks, SAGA reduces task completion time by 1.64x (geometric mean, p < 0.001) over vLLM v0.15.1 with prefix caching and affinity routing, while improving GPU memory utilization by 1.22x and achieving 99.2% SLO attainment under multi-tenant interference. These latency gains come at a quantified cost: approximately 30% lower peak throughput than throughput-optimal batch scheduling, a tradeoff appropriate for the latency-sensitive interactive deployments that dominate compound AI usage. Our results demonstrate that workflow-aware scheduling is essential for efficient compound AI serving.

📄 PDF Abstract BibTeX arXiv:2605.00528

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Atomic-to-Compositional Generalization for Mobile Agents with A New Benchmark and Scheduling System

2025-06-10 · Yuan Guo, Tingjia Miao, Zheng Wu, Pengzhou Cheng 외

Autonomous agents powered by multimodal large language models have been developed to facilitate task execution on mobile devices. However, prior work has predominantly focused on atomic tasks -- such as shot-chain execut…

Scheduling

SimuHome: A Temporal- and Environment-Aware Benchmark for Smart Home LLM Agents

2025-09-29 · Gyuhyeon Seo, Jungwoo Yang, Junseong Pyo, Nalim Kim 외 arxiv

We introduce $\textbf{SimuHome}$, a high-fidelity smart home simulator and a benchmark of 600 episodes for LLM-based smart home agents. Existing smart home benchmarks treat the home as a static system, neither simulating…

A Task-Driven and Quality-Assured Agent Framework for SAR Data Generation

2026-06-27 · Xuanting Wu, Fan Zhanga, Fei Ma, Ling Guan 외 arxiv

Synthetic aperture radar (SAR) data augmentation is important for improving the generalization of data-driven SAR interpretation models, yet practical augmentation workflows are often hindered by heterogeneous dataset fo…

Data Augmentation

SagaLLM: Context Management, Validation, and Transaction Guarantees for Multi-Agent LLM Planning

2025-03-15 · Edward Y. Chang, Longling Geng

Recent LLM-based agent frameworks have demonstrated impressive capabilities in task delegation and workflow orchestration, but face significant challenges in maintaining context awareness and ensuring planning consistenc…

Decision MakingManagement

Throughput-Optimal Scheduling Algorithms for LLM Inference and AI Agents

2025-04-10 · Yueying Li, Jim Dai, Tianyi Peng

As demand for Large Language Models (LLMs) and AI agents rapidly grows, optimizing systems for efficient LLM inference becomes critical. While significant efforts have focused on system-level engineering, little is explo…

AI AgentLarge Language ModelScheduling