paper-with-me

홈 › Papers

GraphCC: A Practical Graph Learning-based Approach to Congestion Control in Datacenters

2023-08-09 · Guillermo Bernárdez, José Suárez-Varela, Xiang Shi, Shihan Xiao, Xiangle Cheng, Pere Barlet-Ros, Albert Cabellos-Aparicio

Congestion Control (CC) plays a fundamental role in optimizing traffic in Data Center Networks (DCN). Currently, DCNs mainly implement two main CC protocols: DCTCP and DCQCN. Both protocols -- and their main variants -- are based on Explicit Congestion Notification (ECN), where intermediate switches mark packets when they detect congestion. The ECN configuration is thus a crucial aspect on the performance of CC protocols. Nowadays, network experts set static ECN parameters carefully selected to optimize the average network performance. However, today's high-speed DCNs experience quick and abrupt changes that severely change the network state (e.g., dynamic traffic workloads, incast events, failures). This leads to under-utilization and sub-optimal performance. This paper presents GraphCC, a novel Machine Learning-based framework for in-network CC optimization. Our distributed solution relies on a novel combination of Multi-agent Reinforcement Learning (MARL) and Graph Neural Networks (GNN), and it is compatible with widely deployed ECN-based CC protocols. GraphCC deploys distributed agents on switches that communicate with their neighbors to cooperate and optimize the global ECN configuration. In our evaluation, we test the performance of GraphCC under a wide variety of scenarios, focusing on the capability of this solution to adapt to new scenarios unseen during training (e.g., new traffic workloads, failures, upgrades). We compare GraphCC with a state-of-the-art MARL-based solution for ECN tuning -- ACC -- and observe that our proposed solution outperforms the state-of-the-art baseline in all of the evaluation scenarios, showing improvements up to $20\%$ in Flow Completion Time as well as significant reductions in buffer occupancy ($38.0-85.7\%$).

📄 PDF Abstract BibTeX arXiv:2308.04905

Code (0)

등록된 구현이 없습니다.

Tasks

Graph LearningMulti-agent Reinforcement Learning

Similar Papers 제목 키워드 기반

Reinforcement Learning for Datacenter Congestion Control

2021-02-18 · Chen Tessler, Yuval Shpigelman, Gal Dalal, Amit Mandelbaum 외

We approach the task of network congestion control in datacenters using Reinforcement Learning (RL). Successful congestion control algorithms can dramatically improve latency and overall network throughput. Until today, …

Network Congestion Controlreinforcement-learningReinforcement LearningReinforcement Learning (RL)

Impact of RoCE Congestion Control Policies on Distributed Training of DNNs

2022-07-22 · Tarannum Khan, Saeed Rashidi, Srinivas Sridharan, Pallavi Shurpali 외

RDMA over Converged Ethernet (RoCE) has gained significant attraction for datacenter networks due to its compatibility with conventional Ethernet-based fabric. However, the RDMA protocol is efficient only on (nearly) los…

Blocking

MOSAIC: A Multi-Objective Optimization Framework for Sustainable Datacenter Management

2023-11-14 · Sirui Qi, Dejan Milojicic, Cullen Bash, Sudeep Pasricha

In recent years, cloud service providers have been building and hosting datacenters across multiple geographical locations to provide robust services. However, the geographical distribution of datacenters introduces grow…

Management

Boundary Control of Traffic Congestion Modeled as a Non-stationary Stochastic Process

2021-03-26 · Xun Liu, Hossein Rastgoftar

In this paper, we introduce a new conservation-based approach to model traffic dynamics, and apply the model predictive control (MPC) approach to control the boundary traffic inflow and outflow, so that the traffic conge…

ManagementModel Predictive Control

Classic Meets Modern: a Pragmatic Learning-Based Congestion Control for the Internet

2020-07-30 · SIGCOMM 2020 7 · Soheil Abbasloo, Chen-Yu Yen, H. Jonathan Chao

These days, taking the revolutionary approach of using clean-slate learning-based designs to completely replace the classic congestion control schemes for the Internet is gaining popularity. However, we argue that curren…

Deep Reinforcement Learning