Machine Learning Fleet Efficiency: Analyzing and Optimizing Large-Scale Google TPU Systems with ML Productivity Goodput
Recent years have seen the emergence of machine learning (ML) workloads deployed in warehouse-scale computing (WSC) settings, also known as ML fleets. As the computational demands placed on ML fleets have increased due to the rise of large models and growing demand for ML applications, it has become increasingly critical to measure and improve the efficiency of such systems. However, there is not yet an established methodology to characterize ML fleet performance and identify potential performance optimizations accordingly. This paper presents a large-scale analysis of an ML fleet based on Google's TPUs, introducing a framework to capture fleet-wide efficiency, systematically evaluate performance characteristics, and identify optimization strategies for the fleet. We begin by defining an ML fleet, outlining its components, and analyzing an example Google ML fleet in production comprising thousands of accelerators running diverse workloads. Our study reveals several critical insights: first, ML fleets extend beyond the hardware layer, with model, data, framework, compiler, and scheduling layers significantly impacting performance; second, the heterogeneous nature of ML fleets poses challenges in characterizing individual workload performance; and third, traditional utilization-based metrics prove insufficient for ML fleet characterization. To address these challenges, we present the "ML Productivity Goodput" (MPG) metric to measure ML fleet efficiency. We show how to leverage this metric to characterize the fleet across the ML system stack. We also present methods to identify and optimize performance bottlenecks using MPG, providing strategies for managing warehouse-scale ML systems in general. Lastly, we demonstrate quantitative evaluations from applying these methods to a real ML fleet for internal-facing Google TPU workloads, where we observed tangible improvements.
Code (0)
등록된 구현이 없습니다.
Tasks
SchedulingSimilar Papers 제목 키워드 기반
Interpretable Machine Learning Models for Predicting and Explaining Vehicle Fuel Consumption Anomalies
Identifying anomalies in the fuel consumption of the vehicles of a fleet is a crucial aspect for optimizing consumption and reduce costs. However, this information alone is insufficient, since fleet operators need to kno…
Anomaly DetectionBIG-bench Machine LearningInterpretable Machine LearningUnsupervised Anomaly DetectionInteger-Clustering Optimization of Hydrogen and Battery EV Fleets Considering DERs
Electrified transportation leads to a tighter integration between transportation and energy distribution systems. In this work, we develop scalable optimization models to co-design hydrogen and battery electric vehicle (…
ClusteringVehicle Fuel Optimization Under Real-World Driving Conditions: An Explainable Artificial Intelligence Approach
Fuel optimization of diesel and petrol vehicles within industrial fleets is critical for mitigating costs and reducing emissions. This objective is achievable by acting on fuel-related factors, such as the driving behavi…
Explainable artificial intelligenceAccelerating Fleet Upgrade Decisions with Machine-Learning Enhanced Optimization
Rental-based business models and increasing sustainability requirements intensify the need for efficient strategies to manage large machine and vehicle fleet renewal and upgrades. Optimized fleet upgrade strategies maxim…
Cost-optimal Fleet Management Strategies for Solar-electric Autonomous Mobility-on-Demand Systems
This paper studies mobility systems that incorporate a substantial solar energy component, generated not only on the ground, but also through solar roofs installed on vehicles, directly covering a portion of their energy…
Autonomous VehiclesManagement