paper-with-me

Papers

Scaling Neural Network Verification with Tensor Parallelism and Fully Sharded Data Parallelism

2026-06-08 · Sergei Vorobyov, Eugene Ilyushin arxiv

Formal neural network verification -- proving that a network satisfies safety properties for *all* inputs in a specified domain -- is bounded in practice by GPU memory: standard implementations of bound-propagation algorithms (IBP, CROWN, $α$-CROWN) require weight and relaxation-coefficient matrices to reside entirely on one accelerator. We adapt two parallelism techniques originally developed for large-scale model training to the auto_LiRPA / $α,β$-CROWN verification framework. Tensor Parallelism (TP) shards both weight and $A$-matrices across GPUs, achieving ${\approx}2\times$ peak-memory reduction at $P{=}2$; soundness is confirmed on VNN-COMP 2022 MNIST-FC benchmarks, though bound tightness degrades with the number of sharded zones due to forced IBP substitution for intermediate bounds inside sharded zones. Fully Sharded Data Parallelism (FSDP) shards only weight matrices with a per-layer AllGather, producing bounds that are bitwise identical to the single-GPU baseline: baseline memory drops by 80--90%, peak memory by 34--39% on wide MLPs. FSDP integrates cleanly with complete verification ($β$-CROWN + Branch-and-Bound) and with convolutional layers (BoundConv); a complete unsat result is obtained for CIFAR-100 ResNet-large (VNN-COMP 2024) under FSDP. Across all experiments the memory bottleneck in $α$-CROWN+BaB mode proves to be per-neuron alpha tensors, not weight matrices, pointing to the key direction for future work.

📄 PDF Abstract BibTeX arXiv:2606.09377

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

The Energy Consumption of Transformer Fine-Tuning: A Roofline-Inspired Scaling Model

2026-06-22 · Mansour Zoubeirou a Mayaki arxiv

Transformer-based models underpin modern natural language processing but incur rapidly growing computational and energy costs. As training scales in both model size and parallelism, accurately predicting energy consumpti…

Placement Semantics for Distributed Deep Learning: A Systematic Framework for Analyzing Parallelism Strategies

2026-01-05 · Deep Pankajbhai Mehta arxiv

Training large language models requires distributing computation across many accelerators, yet practitioners select parallelism strategies (data, tensor, pipeline, ZeRO) through trial and error because no unified systema…

On Optimizing the Communication of Model Parallelism

2022-11-10 · Yonghao Zhuang, Hexu Zhao, Lianmin Zheng, Zhuohan Li 외

We study a novel and important communication pattern in large-scale model-parallel deep learning (DL), which we call cross-mesh resharding. This pattern emerges when the two paradigms of model parallelism - intra-operato…

model

veScale: Consistent and Efficient Tensor Programming with Eager-Mode SPMD

2025-09-05 · Youjie Li, Cheng Wan, Zhiqi Lin, Hongyu Zhu 외 arxiv

Large Language Models (LLMs) have scaled rapidly in size and complexity, requiring increasingly intricate parallelism for distributed training, such as 3D parallelism. This sophistication motivates a shift toward simpler…

Breadth-First Pipeline Parallelism

2022-11-11 · Joel Lamy-Poirier

We introduce Breadth-First Pipeline Parallelism, a novel training schedule which optimizes the combination of pipeline and data parallelism. Breadth-First Pipeline Parallelism lowers training time, cost and memory usage …

GPU