Joint Model Assignment and Resource Allocation for Cost-Effective Mobile Generative Services
Artificial Intelligence Generated Content (AIGC) services can efficiently satisfy user-specified content creation demands, but the high computational requirements pose various challenges to supporting mobile users at scale. In this paper, we present our design of an edge-enabled AIGC service provisioning system to properly assign computing tasks of generative models to edge servers, thereby improving overall user experience and reducing content generation latency. Specifically, once the edge server receives user requested task prompts, it dynamically assigns appropriate models and allocates computing resources based on features of each category of prompts. The generated contents are then delivered to users. The key to this system is a proposed probabilistic model assignment approach, which estimates the quality score of generated contents for each prompt based on category labels. Next, we introduce a heuristic algorithm that enables adaptive configuration of both generation steps and resource allocation, according to the various task requests received by each generative model on the edge.Simulation results demonstrate that the designed system can effectively enhance the quality of generated content by up to 4.7% while reducing response delay by up to 39.1% compared to benchmarks.
Code (0)
등록된 구현이 없습니다.
Methods 이 논문이 사용한 방법론
Similar Papers 제목 키워드 기반
User Assignment and Resource Allocation for Hierarchical Federated Learning over Wireless Networks
The large population of wireless users is a key driver of data-crowdsourced Machine Learning (ML). However, data privacy remains a significant concern. Federated Learning (FL) encourages data sharing in ML without requir…
Combinatorial OptimizationCPUFederated LearningResearch on Resource Allocation for Efficient Federated Learning
As a promising solution to achieve efficient learning among isolated data owners and solve data privacy issues, federated learning is receiving wide attention. Using the edge server as an intermediary can effectively col…
Edge-computingFederated LearningSubcarrier Assignment and Power Allocation for SCMA Energy Efficiency
In this paper we propose resource allocation algorithm for uplink sparse code multiple access (SCMA) networks to maximize the energy efficiency (EE). Due to the joint optimization of factor graph matrix and power allocat…
Robust Batch-Level Query Routing for Large Language Models under Cost and Capacity Constraints
We study the problem of routing queries to large language models (LLMs) under cost, GPU resources, and concurrency constraints. Prior per-query routing methods often fail to control batch-level cost, especially under non…
Joint Radio Resource Allocation and Cooperative Caching in PD-NOMA-Based HetNets
In this paper, we propose a novel joint resource allocation and cooperative caching scheme for power-domain non-orthogonal multiple access (PD-NOMA)-based heterogeneous networks (HetNets). In our scheme, the requested co…
Management