paper-with-me

홈 › Papers

Online detection of failures generated by storage simulator

2021-01-18 · Kenenbek Arzymatov, Mikhail Hushchyn, Andrey Sapronov, Vladislav Belavin, Leonid Gremyachikh, Maksim Karpov, Andrey Ustyuzhanin

Modern large-scale data-farms consist of hundreds of thousands of storage devices that span distributed infrastructure. Devices used in modern data centers (such as controllers, links, SSD- and HDD-disks) can fail due to hardware as well as software problems. Such failures or anomalies can be detected by monitoring the activity of components using machine learning techniques. In order to use these techniques, researchers need plenty of historical data of devices in normal and failure mode for training algorithms. In this work, we challenge two problems: 1) lack of storage data in the methods above by creating a simulator and 2) applying existing online algorithms that can faster detect a failure occurred in one of the components. We created a Go-based (golang) package for simulating the behavior of modern storage infrastructure. The software is based on the discrete-event modeling paradigm and captures the structure and dynamics of high-level storage system building blocks. The package's flexible structure allows us to create a model of a real-world storage system with a configurable number of components. The primary area of interest is exploring the storage machine's behavior under stress testing or exploitation in the medium- or long-term for observing failures of its components. To discover failures in the time series distribution generated by the simulator, we modified a change point detection algorithm that works in online mode. The goal of the change-point detection is to discover differences in time series distribution. This work describes an approach for failure detection in time series data based on direct density ratio estimation via binary classifiers.

📄 PDF Abstract BibTeX arXiv:2101.07100

Code (0)

등록된 구현이 없습니다.

Tasks

Change Point DetectionDensity Ratio EstimationTime SeriesTime Series Analysis

Similar Papers 제목 키워드 기반

DNA Storage Error Simulator: A Tool for Simulating Errors in Synthesis, Storage, PCR and Sequencing

2022-05-28 · Jamie J. Alnasir, Thomas Heinis, Louis Carteron

DNA has many valuable characteristics that make it suitable for a long-term storage medium, in particular its durability and high information density. DNA can be stored safely for hundreds of years with virtually no degr…

Recover: A Neuro-Symbolic Framework for Failure Detection and Recovery

2024-03-31 · Cristina Cornelio, Mohammed Diab

Recognizing failures during task execution and implementing recovery procedures is challenging in robotics. Traditional approaches rely on the availability of extensive data or a tight set of constraints, while more rece…

Finding Failures in High-Fidelity Simulation using Adaptive Stress Testing and the Backward Algorithm

2021-07-27 · Mark Koren, Ahmed Nassar, Mykel J. Kochenderfer

Validating the safety of autonomous systems generally requires the use of high-fidelity simulators that adequately capture the variability of real-world scenarios. However, it is generally not feasible to exhaustively se…

Autonomous VehiclesDeep Reinforcement Learningreinforcement-learningReinforcement Learning (RL)

Code-as-Monitor: Constraint-aware Visual Programming for Reactive and Proactive Robotic Failure Detection

2024-12-05 · CVPR 2025 1 · Enshen Zhou, Qi Su, Cheng Chi, Zhizheng Zhang 외

Automatic detection and prevention of open-set failures are crucial in closed-loop robotic systems. Recent studies often struggle to simultaneously identify unexpected failures reactively after they occur and prevent for…

Language Modelling

Closed-Loop CO2 Storage Control With History-Based Reinforcement Learning and Latent Model-Based Adaptation

2026-05-04 · Sofianos Panagiotis Fotias, Vassilis Gaganis arxiv

Closed-loop management of geological CO2 storage requires control policies that adapt to uncertain reservoir behavior while relying on observations that are realistically available during operation. This work formulates …

Reinforcement Learning