Unsupervised KPIs-Based Clustering of Jobs in HPC Data Centers
Performance analysis is an essential task in High-Performance Computing (HPC) systems and it is applied for different purposes such as anomaly detection, optimal resource allocation, and budget planning. HPC monitoring tasks generate a huge number of Key Performance Indicators (KPIs) to supervise the status of the jobs running in these systems. KPIs give data about CPU usage, memory usage, network (interface) traffic, or other sensors that monitor the hardware. Analyzing this data, it is possible to obtain insightful information about running jobs, such as their characteristics, performance, and failures. The main contribution in this paper is to identify which metric/s (KPIs) is/are the most appropriate to identify/classify different types of jobs according to their behavior in the HPC system. With this aim, we have applied different clustering techniques (partition and hierarchical clustering algorithms) using a real dataset from the Galician Computation Center (CESGA). We have concluded that (i) those metrics (KPIs) related to the Network (interface) traffic monitoring provide the best cohesion and separation to cluster HPC jobs, and (ii) hierarchical clustering algorithms are the most suitable for this task. Our approach was validated using a different real dataset from the same HPC center.
Code (0)
등록된 구현이 없습니다.
Tasks
Anomaly DetectionClusteringCPUSimilar Papers 제목 키워드 기반
Machine Learning Based Prediction and Classification of Computational Jobs in Cloud Computing Centers
With the rapid growth of the data volume and the fast increasing of the computational model complexity in the scenario of cloud computing, it becomes an important topic that how to handle users' requests by scheduling co…
BIG-bench Machine LearningCloud ComputingClusteringGeneral Classification+1KPIs-Based Clustering and Visualization of HPC jobs: a Feature Reduction Approach
High-Performance Computing (HPC) systems need to be constantly monitored to ensure their stability. The monitoring systems collect a tremendous amount of data about different parameters or Key Performance Indicators (KPI…
ClusteringCPUManagementTime SeriesUrban Spatial Structure and the Potential for Vehicle Miles Traveled Reduction: The Effects of Accessibility to Jobs within and beyond Employment Sub-centers
This research examines the relationship between urban polycentric spatial structure and driving. We identified 46 employment sub-centers in the Los Angeles Combined Statistical Area and calculated access to jobs that are…
Seed-Point Detection of Clumped Convex Objects by Short-Range Attractive Long-Range Repulsive Particle Clustering
Locating the center of convex objects is important in both image processing and unsupervised machine learning/data clustering fields. The automated analysis of biological images uses both of these fields for locating cel…
ClusteringSequence-to-sequence models for workload interference
Co-scheduling of jobs in data-centers is a challenging scenario, where jobs can compete for resources yielding to severe slowdowns or failed executions. Efficient job placement on environments where resources are shared …
BIG-bench Machine LearningCPUScheduling