paper-with-me

홈 › Papers

ScaleAcross Explorer: Exploring Communication Optimization for Scale-Across AI Model Training

2026-05-23 · Minghao Li, Alicia Golden, Samuel Hsia, Michael Kuchnik, Adi Gangidi, Xu Zhang, Ashmitha Jeevaraj Shetty, Zachary DeVito, Weiwei Chu, Dong He, Haoci Zhang, Yuchen Hao, Ruoming Pang, James Hongyi Zeng, Ying Zhang, Minlan Yu, Carole-Jean Wu arxiv

The rapid scaling of large language model training requires distributing GPU resources across multiple data center buildings and regions. We refer to such paradigm as "scale-across" training. As infrastructure expands, the system design space becomes increasingly intricate, encompassing new model architectures, hardware heterogeneity, and evolving communication patterns. Drawing from Meta's production experience, we highlight the complexities of deploying training jobs across a few data centers housing hundreds of thousands of GPUs. To accelerate exploration of the large design space and to enable efficient training for frontier model development, we conduct in-depth characterization of three key design dimensions: parallelism placement, parallelism scheduling, and network layer technologies. We then propose ScaleAcross Explorer, an optimizer that considers the interplay of design dimensions and holistically optimizes scale-across training. Testbed experiments and simulations demonstrate up to 64.62% training speedups over production configuration and up to 37.59% training speedups over the state-of-the-art baseline across a wide range of design points.

📄 PDF Abstract BibTeX arXiv:2605.24326

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

ChannelExplorer: Exploring Class Separability Through Activation Channel Visualization

2025-05-06 · Md Rahat-uz- Zaman, Bei Wang, Paul Rosen

Deep neural networks (DNNs) achieve state-of-the-art performance in many vision tasks, yet understanding their internal behavior remains challenging, particularly how different layers and activation channels contribute t…

Exploring the neighbor graph to improve distributional thesauri (Explorer le graphe de voisinage pour am\'eliorer les th\'esaurus distributionnels) [in French]

2014-07-01 · JEPTALNRECITAL 2014 7 · Vincent Claveau, Ewa Kijak, Olivier Ferret
Information Retrieval

OVD-Explorer: Optimism Should Not Be the Sole Pursuit of Exploration in Noisy Environments

2023-12-19 · Jinyi Liu, Zhi Wang, Yan Zheng, Jianye Hao 외

In reinforcement learning, the optimism in the face of uncertainty (OFU) is a mainstream principle for directing exploration towards less explored areas, characterized by higher uncertainty. However, in the presence of e…

continuous-controlContinuous ControlMuJoCo

OVD-Explorer: A General Information-theoretic Exploration Approach for Reinforcement Learning

2021-09-29 · Jinyi Liu, Zhi Wang, Yan Zheng, Jianye Hao 외

Many exploration strategies are built upon the optimism in the face of the uncertainty (OFU) principle for reinforcement learning. However, without considering the aleatoric uncertainty, existing methods may over-explore…

MuJoCoreinforcement-learningReinforcement Learning (RL)

Agentic RAG with Knowledge Graphs for Complex Multi-Hop Reasoning in Real-World Applications

2025-07-22 · Jean Lelong, Adnane Errazine, Annabelle Blangero arxiv

Conventional Retrieval-Augmented Generation (RAG) systems enhance Large Language Models (LLMs) but often fall short on complex queries, delivering limited, extractive answers and struggling with multiple targeted retriev…

Knowledge Graphs