paper-with-me

Papers

Scalability Optimization in Cloud-Based AI Inference Services: Strategies for Real-Time Load Balancing and Automated Scaling

2025-04-16 · Yihong Jin, Ze Yang

The rapid expansion of AI inference services in the cloud necessitates a robust scalability solution to manage dynamic workloads and maintain high performance. This study proposes a comprehensive scalability optimization framework for cloud AI inference services, focusing on real-time load balancing and autoscaling strategies. The proposed model is a hybrid approach that combines reinforcement learning for adaptive load distribution and deep neural networks for accurate demand forecasting. This multi-layered approach enables the system to anticipate workload fluctuations and proactively adjust resources, ensuring maximum resource utilisation and minimising latency. Furthermore, the incorporation of a decentralised decision-making process within the model serves to enhance fault tolerance and reduce response time in scaling operations. Experimental results demonstrate that the proposed model enhances load balancing efficiency by 35\ and reduces response delay by 28\, thereby exhibiting a substantial optimization effect in comparison with conventional scalability solutions.

📄 PDF Abstract BibTeX arXiv:2504.15296

Code (0)

등록된 구현이 없습니다.

Tasks

Decision MakingDemand Forecasting

Similar Papers 제목 키워드 기반

Deploying Foundation Model Powered Agent Services: A Survey

2024-12-18 · Wenchao Xu, Jinyu Chen, Peirong Zheng, Xiaoquan Yi 외

Foundation model (FM) powered agent services are regarded as a promising solution to develop intelligent and personalized applications for advancing toward Artificial General Intelligence (AGI). To achieve high reliabili…

modelModel CompressionSurveyToken Reduction

Privacy-Aware Joint DNN Model Deployment and Partitioning Optimization for Collaborative Edge Inference Services

2025-02-22 · Zhipeng Cheng, Xiaoyu Xia, Hong Wang, Minghui LiWang 외

Edge inference (EI) has emerged as a promising paradigm to address the growing limitations of cloud-based Deep Neural Network (DNN) inference services, such as high response latency, limited scalability, and severe data …

Stochastic Optimization

Feature Selection using the concept of Peafowl Mating in IDS

2024-02-03 · Partha Ghosh, Joy Sharma, Nilesh Pandey

Cloud computing has high applicability as an Internet based service that relies on sharing computing resources. Cloud computing provides services that are Infrastructure based, Platform based and Software based. The popu…

Cloud Computingfeature selectionIntrusion Detection

Edge-First Language Model Inference: Models, Metrics, and Tradeoffs

2025-05-22 · SiYoung Jang, Roberto Morabito

The widespread adoption of Language Models (LMs) across industries is driving interest in deploying these services across the computing continuum, from the cloud to the network edge. This shift aims to reduce costs, lowe…

BenchmarkingLanguage ModelingLanguage ModellingModel Compression

JointDNN: An Efficient Training and Inference Engine for Intelligent Mobile Cloud Computing Services

2018-01-25 · Amir Erfan Eshratifar, Mohammad Saeed Abrishami, Massoud Pedram

Deep learning models are being deployed in many mobile intelligent applications. End-side services, such as intelligent personal assistants, autonomous cars, and smart home services often employ either simple local model…

Cloud Computing