paper-with-me

홈 › Papers

Ensemble Method for System Failure Detection Using Large-Scale Telemetry Data

2024-06-07 · Priyanka Mudgal, Rita H. Wouhaybi

The growing reliance on computer systems, particularly personal computers (PCs), necessitates heightened reliability to uphold user satisfaction. This research paper presents an in-depth analysis of extensive system telemetry data, proposing an ensemble methodology for detecting system failures. Our approach entails scrutinizing various parameters of system metrics, encompassing CPU utilization, memory utilization, disk activity, CPU temperature, and pertinent system metadata such as system age, usage patterns, core count, and processor type. The proposed ensemble technique integrates a diverse set of algorithms, including Long Short-Term Memory (LSTM) networks, isolation forests, one-class support vector machines (OCSVM), and local outlier factors (LOF), to effectively discern system failures. Specifically, the LSTM network with other machine learning techniques is trained on Intel Computing Improvement Program (ICIP) telemetry software data to distinguish between normal and failed system patterns. Experimental evaluations demonstrate the remarkable efficacy of our models, achieving a notable detection rate in identifying system failures. Our research contributes to advancing the field of system reliability and offers practical insights for enhancing user experience in computing environments.

📄 PDF Abstract BibTeX arXiv:2407.00048

Code (0)

등록된 구현이 없습니다.

Tasks

CPU

Methods 이 논문이 사용한 방법론

Uphold 설명 없음
SET Dynamic Sparse Training method where weight mask is updated randomly periodically
Sigmoid Activation 설명 없음
Tanh Activation 설명 없음
LSTM An LSTM is a type of recurrent neural network that addresses the vanishing gradient problem in vanilla…

Similar Papers 제목 키워드 기반

Hybrid Autoencoder-Based Framework for Early Fault Detection in Wind Turbines

2025-10-16 · Rekha R Nair, Tina Babu, Alavikunhu Panthakkan, Balamurugan Balusamy 외 arxiv

Wind turbine reliability is critical to the growing renewable energy sector, where early fault detection significantly reduces downtime and maintenance costs. This paper introduces a novel ensemble-based deep learning fr…

Unsupervised Anomaly DetectionFeature Engineering

Lost in the Folds: When Cross-Validation Is Not a Deep Ensemble for Uncertainty Estimation

2026-05-18 · Tristan Kirscher, Markus Bujotzek, Yannick Kirchhoff, Maximilian Rokuss 외 arxiv

Ensemble disagreement is widely used as a proxy for epistemic uncertainty in medical image segmentation. In practice, many studies form ensembles via K-fold cross-validation (CV), yet refer to them as ``deep ensembles'' …

Medical Image Segmentation

Ensemble neuroevolution based approach for multivariate time series anomaly detection

2021-08-08 · Kamil Faber, Dominik Żurek, Marcin Pietroń, Kamil Piętak

Multivariate time series anomaly detection is a very common problem in the field of failure prevention. Fast prevention means lower repair costs and losses. The amount of sensors in novel industry systems makes the anoma…

Anomaly DetectionDeep LearningTime SeriesTime Series Analysis+1

Runtime Failure Hunting for Physics Engine Based Software Systems: How Far Can We Go?

2025-07-29 · Shuqing Li, Qiang Chen, Xiaoxue Ren, Michael R. Lyu arxiv

Physics Engines (PEs) are fundamental software frameworks that simulate physical interactions in applications ranging from entertainment to safety-critical systems. Despite their importance, PEs suffer from physics failu…

Autonomous Vehicles

PWPAE: An Ensemble Framework for Concept Drift Adaptation in IoT Data Streams

2021-09-10 · Li Yang, Dimitrios Michael Manias, Abdallah Shami

As the number of Internet of Things (IoT) devices and systems have surged, IoT data analytics techniques have been developed to detect malicious cyber-attacks and secure IoT systems; however, concept drift issues often o…

Anomaly Detection