paper-with-me

Papers

Rearchitecting Datacenter Lifecycle for AI: A TCO-Driven Framework

2025-09-30 · Jovan Stojkovic, Chaojie Zhang, Íñigo Goiri, Ricardo Bianchini arxiv

The rapid rise of large language models (LLMs) has been driving an enormous demand for AI inference infrastructure, mainly powered by high-end GPUs. While these accelerators offer immense computational power, they incur high capital and operational costs due to frequent upgrades, dense power consumption, and cooling demands, making total cost of ownership (TCO) for AI datacenters a critical concern for cloud providers. Unfortunately, traditional datacenter lifecycle management (designed for general-purpose workloads) struggles to keep pace with AI's fast-evolving models, rising resource needs, and diverse hardware profiles. In this paper, we rethink the AI datacenter lifecycle scheme across three stages: building, hardware refresh, and operation. We show how design choices in power, cooling, and networking provisioning impact long-term TCO. We also explore refresh strategies aligned with hardware trends. Finally, we use operation software optimizations to reduce cost. While these optimizations at each stage yield benefits, unlocking the full potential requires rethinking the entire lifecycle. Thus, we present a holistic lifecycle management framework that coordinates and co-optimizes decisions across all three stages, accounting for workload dynamics, hardware evolution, and system aging. Our system reduces the TCO by up to 40\% over traditional approaches. Using our framework we provide guidelines on how to manage AI datacenter lifecycle for the future.

📄 PDF Abstract BibTeX arXiv:2509.26534

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

MARLIN: Multi-Agent Game-Theoretic Reinforcement Learning for Sustainable LLM Inference in Cloud Datacenters

2026-05-13 · H. Moore, S. Qi, D. Milojicic, C. Bash 외 arxiv

Large Language Models (LLMs) have become increasingly prevalent in cloud-based platforms, propelled by the introduction of AI-based consumer and enterprise services. LLM inference requests in particular account for up to…

Reinforcement Learning

OpenG2G: A Simulation Platform for AI Datacenter-Grid Runtime Coordination

2026-05-06 · Jae-Won Chung, Zhirui Liang, Yanyong Mao, Jiasi Chen 외 arxiv

AI's growing compute demand and new datacenter buildouts present major capacity and reliability challenges for the electricity grid, leading to multi-year interconnection delays for new datacenters and bottlenecking AI g…

POLCA: Power Oversubscription in LLM Cloud Providers

2023-08-24 · Pratyush Patel, Esha Choukse, Chaojie Zhang, Íñigo Goiri 외

Recent innovation in large language models (LLMs), and their myriad use-cases have rapidly driven up the compute capacity demand for datacenter GPUs. Several cloud providers and other enterprises have made substantial pl…

GPU

MOSAIC: A Multi-Objective Optimization Framework for Sustainable Datacenter Management

2023-11-14 · Sirui Qi, Dejan Milojicic, Cullen Bash, Sudeep Pasricha

In recent years, cloud service providers have been building and hosting datacenters across multiple geographical locations to provide robust services. However, the geographical distribution of datacenters introduces grow…

Management

Beyond Efficiency: Scaling AI Sustainably

2024-06-08 · Carole-Jean Wu, Bilge Acun, Ramya Raghavendra, Kim Hazelwood

Barroso's seminal contributions in energy-proportional warehouse-scale computing launched an era where modern datacenters have become more energy efficient and cost effective than ever before. At the same time, modern AI…

Deep Learning