paper-with-me

홈 › Papers

Distributed Training under Packet Loss

2025-07-02 · Erez Weintraub, Ron Banner, Ariel Orda arxiv

State-of-the-art language and vision models are routinely trained across thousands of GPUs, often spanning multiple data-centers, yet today's distributed frameworks still assume reliable connections (e.g., InfiniBand or RoCE). The resulting acknowledgment traffic and retransmissions inflate tail latencies and limit scalability. Leveraging unreliable connections will reduce latency but may sacrifice model accuracy and convergence once packets are dropped. A principled, end-to-end solution that preserves accuracy and convergence guarantees under genuine packet loss has previously been missing. We address this critical gap by introducing a novel distributed training framework capable of operating over unreliable connections, offering unbiased gradient aggregation and bounded parameter drift without modifying model code or optimizers. The key insight is a two-stage defense against missing messages: (i) Unbiased gradient aggregation: each worker reconstructs a consistent gradient estimate from whatever packets arrive, guaranteeing expectation-level correctness; and (ii) Bounded-drift parameter broadcasts: we prove the inter-worker model discrepancy remains O(1) even after arbitrarily many iterations, preventing the unbounded divergence typical of asynchronous setups. Analytical bounds are matched by experiments on the LLAMA2 7B model with 64 GPUs: tolerating 10% random packet loss yields at most 0.8% perplexity change. This work bridges the gap between communication-efficient datacenter protocols and the accuracy and generalization guarantees demanded by modern large-model training, enabling robust, high-throughput learning on commodity or wide-area networks.

📄 PDF Abstract BibTeX arXiv:2507.07114

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Packet-Loss-Tolerant Split Inference for Delay-Sensitive Deep Learning in Lossy Wireless Networks

2021-04-28 · Sohei Itahara, Takayuki Nishio, Koji Yamamoto

The distributed inference framework is an emerging technology for real-time applications empowered by cutting-edge deep machine learning (ML) on resource-constrained Internet of things (IoT) devices. In distributed infer…

Distributed fusion filter over lossy wireless sensor networks with the presence of non-Gaussian noise

2023-07-04 · Jiacheng He, Bei Peng, Zhenyu Feng, Xuemei Mao 외

The information transmission between nodes in a wireless sensor networks (WSNs) often causes packet loss due to denial-of-service (DoS) attack, energy limitations, and environmental factors, and the information that is s…

State Estimation

Machine Learning-Based Distributed Authentication of UWAN Nodes with Limited Shared Information

2022-08-19 · Francesco Ardizzon, Roee Diamant, Paolo Casari, Stefano Tomasin

We propose a technique to authenticate received packets in underwater acoustic networks based on the physical layer features of the underwater acoustic channel (UWAC). Several sensors a) locally estimate features (e.g., …

Distributed Decisions on Optimal Load Balancing in Loss Networks

2023-07-10 · Qiong Liu, Chehao Wang, Ce Zheng

When multiple users share a common link in direct transmission, packet loss and network collision may occur due to the simultaneous arrival of traffics at the source node. To tackle this problem, users may resort to an i…

Loss Tolerant Federated Learning

2021-05-08 · Pengyuan Zhou, Pei Fang, Pan Hui

Federated learning has attracted attention in recent years for collaboratively training data on distributed devices with privacy-preservation. The limited network capacity of mobile and IoT devices has been seen as one o…

FairnessFederated Learning