KPIs-Based Clustering and Visualization of HPC jobs: a Feature Reduction Approach
High-Performance Computing (HPC) systems need to be constantly monitored to ensure their stability. The monitoring systems collect a tremendous amount of data about different parameters or Key Performance Indicators (KPIs), such as resource usage, IO waiting time, etc. A proper analysis of this data, usually stored as time series, can provide insight in choosing the right management strategies as well as the early detection of issues. In this paper, we introduce a methodology to cluster HPC jobs according to their KPI indicators. Our approach reduces the inherent high dimensionality of the collected data by applying two techniques to the time series: literature-based and variance-based feature extraction. We also define a procedure to visualize the obtained clusters by combining the two previous approaches and the Principal Component Analysis (PCA). Finally, we have validated our contributions on a real data set to conclude that those KPIs related to CPU usage provide the best cohesion and separation for clustering analysis and the good results of our visualization methodology.
Code (0)
등록된 구현이 없습니다.
Tasks
ClusteringCPUManagementTime SeriesMethods 이 논문이 사용한 방법론
Similar Papers 제목 키워드 기반
Unsupervised KPIs-Based Clustering of Jobs in HPC Data Centers
Performance analysis is an essential task in High-Performance Computing (HPC) systems and it is applied for different purposes such as anomaly detection, optimal resource allocation, and budget planning. HPC monitoring t…
Anomaly DetectionClusteringCPUOptimizing Unlicensed Band Spectrum Sharing With Subspace-Based Pareto Tracing
To meet the ever-growing demands of data throughput for forthcoming and deployed wireless networks, new wireless technologies like Long-Term Evolution License-Assisted Access (LTE-LAA) operate in shared and unlicensed ba…
Dimensionality ReductionMulti-Criteria Radio Spectrum Sharing With Subspace-Based Pareto Tracing
Radio spectrum is a high-demand finite resource. To meet growing demands of data throughput for forthcoming and deployed wireless networks, new wireless technologies must operate in shared spectrum over unlicensed bands …
Dimensionality ReductionConsistent Representation Learning for High Dimensional Data Analysis
High dimensional data analysis for exploration and discovery includes three fundamental tasks: dimensionality reduction, clustering, and visualization. When the three associated tasks are done separately, as is often the…
ClusteringDimensionality ReductionRepresentation LearningVocal Bursts Intensity PredictionDeep Temporal Clustering: Fully unsupervised learning of time-domain features
Unsupervised learning of timeseries data is a challenging problem in machine learning. Here, we propose a novel algorithm, Deep Temporal Clustering (DTC), a fully unsupervised method, to naturally integrate dimensionali…
ClusteringDimensionality Reduction