paper-with-me

Papers

TrainMover: An Interruption-Resilient and Reliable ML Training Runtime

2024-12-17 · ChonLam Lao, Minlan Yu, Aditya Akella, Jiamin Cao, Yu Guan, Pengcheng Zhang, Zhilong Zheng, Yichi Xu, Ennan Zhai, Dennis Cai, Jiaqi Gao

Large-scale ML training jobs are frequently interrupted by hardware and software anomalies, failures, and management events. Existing solutions like checkpointing or runtime reconfiguration suffer from long downtimes, degraded performance, or undesired changes to training strategies. We present TrainMover, a resilient runtime that leverages standby machines to handle interruptions with minimal downtime and zero memory overhead. To achieve these goals, TrainMover introduces two key techniques: two-phase, delta-based communication group setups and communication-free sandboxed shadow iterations. Our evaluation shows that TrainMover consistently achieves second-level downtime across all evaluated models during migration, maintaining 99\% training efficiency during periodic 10-minute rebalancing. We also demonstrate the effectiveness of TrainMover in handling various interruptions.

📄 PDF Abstract BibTeX arXiv:2412.12636

Code (0)

등록된 구현이 없습니다.

Tasks

ManagementScheduling

Similar Papers 제목 키워드 기반

Resilient Time-Varying Output Formation Tracking of Linear Multi-Agent Systems Against Unbounded FDI Sensor Attacks and Unreliable Digraphs

2021-05-06 · Zhi Feng, Guoqiang Hu

One salient feature of cooperative formation tracking is its distributed nature that relies on localized control and information sharing over a sparse communication network. That is, a distributed control manner could be…

Resilient Microgrid Formation Considering Communication Interruptions

2024-03-02 · Jian Zhong, Chen Chen, Young-Jin Kim, Yuxiong Huang 외

Distribution system (DS) communication failures following extreme events often degrade monitoring and control functions, thus preventing the acquisition of complete global DS component state information, on which existin…

CodeNet: Training Large Scale Neural Networks in Presence of Soft-Errors

2019-03-04 · Sanghamitra Dutta, Ziqian Bai, Tze Meng Low, Pulkit Grover

This work proposes the first strategy to make distributed training of neural networks resilient to computing errors, a problem that has remained unsolved despite being first posed in 1956 by von Neumann. He also speculat…

Straggler-Resilient Federated Learning: Leveraging the Interplay Between Statistical Accuracy and System Heterogeneity

2020-12-28 · Amirhossein Reisizadeh, Isidoros Tziotis, Hamed Hassani, Aryan Mokhtari 외

Federated Learning is a novel paradigm that involves learning from data samples distributed across a large network of clients while the data remains local. It is, however, known that federated learning is prone to multip…

Federated Learning

Resilient Identification of Distribution Network Topology

2020-11-16 · Mohammad Jafarian, Alireza Soroudi, Andrew Keane

Network topology identification (TI) is an essential function for distributed energy resources management systems (DERMS) to organize and operate widespread distributed energy resources (DERs). In this paper, discriminan…

Management