paper-with-me

Papers

Linearized Data Center Workload and Cooling Management

2023-04-10 · Somayye Rostami, Douglas G. Down, George Karakostas

With the current high levels of energy consumption of data centers, reducing power consumption by even a small percentage is beneficial. We propose a framework for thermal-aware workload distribution in a data center to reduce cooling power consumption. The framework includes linearization of the general optimization problem and proposing a heuristic to approximate the solution for the resulting Integer Linear Programming (ILP) problems. We first define a general nonlinear power optimization problem including several cooling parameters, heat recirculation effects, and constraints on server temperatures. We propose to study a linearized version of the problem, which is easier to analyze. As an energy saving scenario and as a proof of concept for our approach, we also consider the possibility that the red-line temperature for idle servers is higher than that for busy servers. For the resulting ILP problem, we propose a heuristic for intelligent rounding of the fractional solution. Through numerical simulations, we compare our heuristics with two baseline algorithms. We also evaluate the performance of the solution of the linearized system on the original system. The results show that the proposed approach can reduce the cooling power consumption by more than 30 percent compared to the case of continuous utilizations and a single red-line temperature.

📄 PDF Abstract BibTeX arXiv:2304.04731

Code (0)

등록된 구현이 없습니다.

Tasks

Management

Similar Papers 제목 키워드 기반

LC-Opt: Benchmarking Reinforcement Learning and Agentic AI for End-to-End Liquid Cooling Optimization in Data Centers

2025-10-31 · Avisek Naug, Antonio Guillen, Vineet Kumar, Scott Greenwood 외 arxiv

Liquid cooling is critical for thermal management in high-density data centers with the rising AI workloads. However, machine learning-based controllers are essential to unlock greater energy efficiency and reliability, …

Reinforcement Learning

Rearchitecting Datacenter Lifecycle for AI: A TCO-Driven Framework

2025-09-30 · Jovan Stojkovic, Chaojie Zhang, Íñigo Goiri, Ricardo Bianchini arxiv

The rapid rise of large language models (LLMs) has been driving an enormous demand for AI inference infrastructure, mainly powered by high-end GPUs. While these accelerators offer immense computational power, they incur …

TAPAS: Thermal- and Power-Aware Scheduling for LLM Inference in Cloud Platforms

2025-01-05 · Jovan Stojkovic, Chaojie Zhang, Íñigo Goiri, Esha Choukse 외

The rising demand for generative large language models (LLMs) poses challenges for thermal and power management in cloud datacenters. Traditional techniques often are inadequate for LLM inference due to the fine-grained,…

GPUQuantizationScheduling

Generative Design for Direct-to-Chip Liquid Cooling for Data Centers

2026-04-13 · Zheng Liu arxiv

Rapid growth in artificial intelligence (AI) workloads is driving up data center power densities, increasing the need for advanced thermal management. Direct-to-chip liquid cooling can remove heat efficiently at the sour…

Hierarchical Multi-Agent Framework for Carbon-Efficient Liquid-Cooled Data Center Clusters

2025-02-12 · Soumyendu Sarkar, Avisek Naug, Antonio Guillen, Vineet Gundecha 외

Reducing the environmental impact of cloud computing requires efficient workload distribution across geographically dispersed Data Center Clusters (DCCs) and simultaneously optimizing liquid and air (HVAC) cooling with t…

Cloud ComputingReinforcement Learning (RL)