paper-with-me

홈 › Papers

Cost-Aware Model Orchestration for LLM-based Systems

2025-11-30 · Daria Smirnova, Hamid Nasiri, Marta Adamska, Zhengxin Yu, Peter Garraghan arxiv

As modern artificial intelligence (AI) systems become more advanced and capable, they can leverage a wide range of tools and models to perform complex tasks. The task of orchestrating these models is increasingly performed by Large Language Models (LLMs) that rely on qualitative descriptions of models for decision-making. However, the descriptions provided to existing LLM-based orchestrators frequently do not reflect true model capabilities and performance characteristics, leading to suboptimal model selection, reduced task accuracy, and increased cost. In this paper, we conduct an empirical analysis of LLM-based orchestration limitations and propose a cost-aware model selection method that accounts for performance-cost trade-offs by incorporating quantitative model performance characteristics within decision-making. Initial experimental results demonstrate that our proposed method increases accuracy by 0.90%-11.92% across various evaluated tasks, achieves up to a 54% energy efficiency improvement, and reduces orchestrator model selection latency from 4.51 s to 7.2 ms.

📄 PDF Abstract BibTeX arXiv:2512.01099

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Training-Free Multimodal Large Language Model Orchestration

2025-08-06 · Tianyu Xie, Yuexiao Ma, Yuhang Wu, Wang Chen 외 arxiv

Building interactive omni-modal assistants often relies on end-to-end multimodal alignment to fuse heterogeneous modalities, which incurs substantial data and compute costs and limits extensibility. We present Training-F…

Learning Latency-Aware Orchestration for Parallel Multi-Agent Systems

2026-01-15 · Xi Shi, Mengxin Zheng, Qian Lou arxiv

Multi-agent systems (MAS) enable complex reasoning by coordinating multiple agents, but often incur high inference latency due to multi-step execution and repeated model invocations, severely limiting their scalability a…

IslandRun: Privacy-Aware Multi-Objective Orchestration for Distributed AI Inference

2025-11-29 · Bala Siva Sai Akhil Malepati arxiv

Modern AI inference faces an irreducible tension: no single computational resource simultaneously maximizes performance, preserves privacy, minimizes cost, and maintains trust. Existing orchestration frameworks optimize …

Federated Learning

When Parallelism Pays Off: Cohesion-Aware Task Partitioning for Multi-Agent Coding

2026-05-31 · Xu Yang, Lunyiu Nie, Ethan Chandra, Stanislav Gannutin 외 arxiv

Multi-agent Large Language Model (LLM) systems offer a way to decompose complex tasks, such as coding, through parallelization and context isolation. However, adding agents in practice introduces inter-agent communicatio…

Community Detectiongraph partitioning

Difficulty-Aware Agentic Orchestration for Query-Specific Multi-Agent Workflows

2025-09-14 · Jinwei Su, Qizhen Lan, Yinghui Xia, Lifan Sun 외 arxiv

Large Language Model (LLM)-based agentic systems have shown strong capabilities across various tasks. However, existing multi-agent frameworks often rely on static or task-level workflows, which either over-process simpl…