ScaleAcross Explorer: Exploring Communication Optimization for Scale-Across AI Model Training
The rapid scaling of large language model training requires distributing GPU resources across multiple data center buildings and regions. We refer to such paradigm as "scale-across" training. As infrastructure expands, the system design space becomes increasingly intricate, encompassing new model architectures, hardware heterogeneity, and evolving communication patterns. Drawing from Meta's production experience, we highlight the complexities of deploying training jobs across a few data centers housing hundreds of thousands of GPUs. To accelerate exploration of the large design space and to enable efficient training for frontier model development, we conduct in-depth characterization of three key design dimensions: parallelism placement, parallelism scheduling, and network layer technologies. We then propose ScaleAcross Explorer, an optimizer that considers the interplay of design dimensions and holistically optimizes scale-across training. Testbed experiments and simulations demonstrate up to 64.62% training speedups over production configuration and up to 37.59% training speedups over the state-of-the-art baseline across a wide range of design points.
Code (0)
등록된 구현이 없습니다.
Similar Papers 제목 키워드 기반
ChannelExplorer: Exploring Class Separability Through Activation Channel Visualization
Deep neural networks (DNNs) achieve state-of-the-art performance in many vision tasks, yet understanding their internal behavior remains challenging, particularly how different layers and activation channels contribute t…
Exploring the neighbor graph to improve distributional thesauri (Explorer le graphe de voisinage pour am\'eliorer les th\'esaurus distributionnels) [in French]
OVD-Explorer: Optimism Should Not Be the Sole Pursuit of Exploration in Noisy Environments
In reinforcement learning, the optimism in the face of uncertainty (OFU) is a mainstream principle for directing exploration towards less explored areas, characterized by higher uncertainty. However, in the presence of e…
continuous-controlContinuous ControlMuJoCoOVD-Explorer: A General Information-theoretic Exploration Approach for Reinforcement Learning
Many exploration strategies are built upon the optimism in the face of the uncertainty (OFU) principle for reinforcement learning. However, without considering the aleatoric uncertainty, existing methods may over-explore…
MuJoCoreinforcement-learningReinforcement Learning (RL)Agentic RAG with Knowledge Graphs for Complex Multi-Hop Reasoning in Real-World Applications
Conventional Retrieval-Augmented Generation (RAG) systems enhance Large Language Models (LLMs) but often fall short on complex queries, delivering limited, extractive answers and struggling with multiple targeted retriev…
Knowledge Graphs