paper-with-me

Papers

LC-Opt: Benchmarking Reinforcement Learning and Agentic AI for End-to-End Liquid Cooling Optimization in Data Centers

2025-10-31 · Avisek Naug, Antonio Guillen, Vineet Kumar, Scott Greenwood, Wesley Brewer, Sahand Ghorbanpour, Ashwin Ramesh Babu, Vineet Gundecha, Ricardo Luna Gutierrez, Soumyendu Sarkar arxiv

Liquid cooling is critical for thermal management in high-density data centers with the rising AI workloads. However, machine learning-based controllers are essential to unlock greater energy efficiency and reliability, promoting sustainability. We present LC-Opt, a Sustainable Liquid Cooling (LC) benchmark environment, for reinforcement learning (RL) control strategies in energy-efficient liquid cooling of high-performance computing (HPC) systems. Built on the baseline of a high-fidelity digital twin of Oak Ridge National Lab's Frontier Supercomputer cooling system, LC-Opt provides detailed Modelica-based end-to-end models spanning site-level cooling towers to data center cabinets and server blade groups. RL agents optimize critical thermal controls like liquid supply temperature, flow rate, and granular valve actuation at the IT cabinet level, as well as cooling tower (CT) setpoints through a Gymnasium interface, with dynamic changes in workloads. This environment creates a multi-objective real-time optimization challenge balancing local thermal regulation and global energy efficiency, and also supports additional components like a heat recovery unit (HRU). We benchmark centralized and decentralized multi-agent RL approaches, demonstrate policy distillation into decision and regression trees for interpretable control, and explore LLM-based methods that explain control actions in natural language through an agentic mesh architecture designed to foster user trust and simplify system management. LC-Opt democratizes access to detailed, customizable liquid cooling models, enabling the ML community, operators, and vendors to develop sustainable data center liquid cooling control solutions.

📄 PDF Abstract BibTeX arXiv:2511.00116

Code (0)

등록된 구현이 없습니다.

Tasks

Reinforcement Learning

Similar Papers 제목 키워드 기반

Liquid Cooling System for a High Power, Medium Frequency, and Medium Voltage Isolated Power Converter

2023-10-05 · Hooman Taghavi, Ahmad El Shafei, Adel Nasiri

Power electronics systems, widely used in various applications such as industrial automation, electric cars, and renewable energy, have the primary function of converting and controlling electrical power to the desired t…

Hierarchical Multi-Agent Framework for Carbon-Efficient Liquid-Cooled Data Center Clusters

2025-02-12 · Soumyendu Sarkar, Avisek Naug, Antonio Guillen, Vineet Gundecha 외

Reducing the environmental impact of cloud computing requires efficient workload distribution across geographically dispersed Data Center Clusters (DCCs) and simultaneously optimizing liquid and air (HVAC) cooling with t…

Cloud ComputingReinforcement Learning (RL)

A Digital Twin Framework for Liquid-cooled Supercomputers as Demonstrated at Exascale

2024-10-07 · Wesley Brewer, Matthias Maiterth, Vineet Kumar, Rafal Wojda 외

We present ExaDigiT, an open-source framework for developing comprehensive digital twins of liquid-cooled supercomputers. It integrates three main modules: (1) a resource allocator and power simulator, (2) a transient th…

A Compact Hybrid Battery Thermal Management System for Enhanced Cooling

2024-12-01 · Zhipeng Lyu, Jinrong Su, Zhe Li, Xiang Li 외

Hybrid battery thermal management systems (HBTMS) combining active liquid cooling and passive phase change materials (PCM) cooling have shown a potential for the thermal management of lithium-ion batteries. However, the …

Management

High Efficiency Polymer based Direct Multi-jet Impingement Cooling Solution for High Power Devices

2023-10-18 · Tiwei Wei

Liquid jet impingement cooling is an efficient cooling technique where the liquid coolant is directly ejected from nozzles on the chip backside resulting in a high cooling efficiency due to the absence of the TIM and the…