paper-with-me

홈 › Papers

HiveMind: OS-Inspired Scheduling for Concurrent LLM Agent Workloads

2026-04-18 · Justice Owusu Agyemang, Jerry John Kponyo, Obed Kwasi Somuah, Elliot Amponsah, Godfred Manu Addo Boakye, Kwame Opuni-Boachie Obour Agyekum arxiv

When multiple LLM coding agents share a rate-limited API endpoint, they exhibit resource contention patterns analogous to unscheduled OS processes competing for CPU, memory, and I/O. In a motivating incident, 3 of 11 parallel agents died from connection resets and HTTP 502 errors - a 27% failure rate - despite the API having sufficient aggregate capacity to serve all 11 sequentially. We present HIVEMIND, a transparent HTTP proxy that applies five OS-inspired scheduling primitives - admission control, rate-limit tracking, AIMD backpressure with circuit breaking, token budget management, and priority queuing - to eliminate the failure modes caused by uncoordinated parallel execution. The proxy requires zero modifications to existing agent code and supports Anthropic, OpenAI, and local model APIs via auto-detected provider profiles. Our evaluation across seven scenarios (5-50 concurrent agents) shows that uncoordinated agents fail at 72-100% rates under contention, while HIVEMIND reduces failures to 0-18% and eliminates 48-100% of wasted compute. An ablation study reveals that transparent retry - not admission control - is the single most critical primitive, but the primitives are most effective in combination. Real-world validation against Ollama confirms that HIVEMIND adds under 3ms of proxy overhead per request. The system is open-source under the MIT license.

📄 PDF Abstract BibTeX arXiv:2604.17111

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

ConEQsA: Concurrent and Asynchronous Embodied Questions Scheduling and Answering

2025-09-15 · Haisheng Wang, Dong Liu, Weiming Zhi arxiv

This paper formulates the Embodied Questions Answering (EQsA) problem, introduces a corresponding benchmark, and proposes an agentic system to tackle the problem. Classical Embodied Question Answering (EQA) is typically …

Question Answering

Towards Understanding, Analyzing, and Optimizing Agentic AI Execution: A CPU-Centric Perspective

2025-11-01 · Ritik Raj, Souvik Kundu, Ishita Vohra, Hong Wang 외 arxiv

Agentic AI serving converts monolithic LLM-based inference to autonomous problem-solvers that can plan, call tools, perform reasoning, and adapt on the fly. Due to diverse task execution need, such serving heavily rely o…

Twill: Scheduling Compound AI Systems on Heterogeneous Mobile Edge Platforms

2025-07-01 · Zain Taufique, Aman Vyas, Antonio Miele, Pasi Liljeberg 외 arxiv

Compound AI (cAI) systems chain multiple AI models to solve complex problems. cAI systems are typically composed of deep neural networks (DNNs), transformers, and large language models (LLMs), exhibiting a high degree of…

Deep Reinforcement Agent for Scheduling in HPC

2021-02-11 · Yuping Fan, Zhiling Lan, Taylor Childers, Paul Rich 외

Cluster scheduler is crucial in high-performance computing (HPC). It determines when and which user jobs should be allocated to available system resources. Existing cluster scheduling heuristics are developed by human ex…

Deep Reinforcement LearningScheduling

Exploring the Dynamic Scheduling Space of Real-Time Generative AI Applications on Emerging Heterogeneous Systems

2025-07-19 · Rachid Karami, Rajeev Patwari, Hyoukjun Kwon, Ashish Sirasao arxiv

The integration of generative AI models, particularly large language models (LLMs), into real-time multi-model AI applications such as video conferencing and gaming is giving rise to a new class of workloads: real-time g…