paper-with-me

Papers

AXLearn: Modular Large Model Training on Heterogeneous Infrastructure

2025-07-07 · Mark Lee, Tom Gunter, Chang Lan, John Peebles, Hanzhi Zhou, Kelvin Zou, Sneha Bangalore, Chung-Cheng Chiu, Nan Du, Xianzhi Du, Philipp Dufter, Ruixuan Hou, Haoshuo Huang, Dongseong Hwang, Xiang Kong, Jinhao Lei, Tao Lei, Meng Li, Li Li, Jiarui Lu, Zhiyun Lu, Yiping Ma, David Qiu, Vivek Rathod, Senyu Tong, Zhucheng Tu, Jianyu Wang, Yongqiang Wang, ZiRui Wang, Floris Weers, Sam Wiseman, Guoli Yin, BoWen Zhang, Xiyou Zhou, Danyang Zhuo, Cheng Leong, Ruoming Pang

We design and implement AXLearn, a production deep learning system that facilitates scalable and high-performance training of large deep learning models. Compared to other state-of-the-art deep learning systems, AXLearn has a unique focus on modularity and support for heterogeneous hardware infrastructure. AXLearn's internal interfaces between software components follow strict encapsulation, allowing different components to be assembled to facilitate rapid model development and experimentation on heterogeneous compute infrastructure. We introduce a novel method of quantifying modularity via Lines-of-Code (LoC)-complexity, which demonstrates how our system maintains constant complexity as we scale the components in the system, compared to linear or quadratic complexity in other systems. This allows integrating features such as Rotary Position Embeddings (RoPE) into AXLearn across hundred of modules with just 10 lines of code, compared to hundreds as required in other systems. At the same time, AXLearn maintains equivalent performance compared to state-of-the-art training systems. Finally, we share our experience in the development and operation of AXLearn.

📄 PDF Abstract BibTeX arXiv:2507.05411

Code (0)

등록된 구현이 없습니다.

Tasks

Deep Learning

Methods 이 논문이 사용한 방법론

Focus 설명 없음

Similar Papers 제목 키워드 기반

HetSeq: Distributed GPU Training on Heterogeneous Infrastructure

2020-09-25 · Yifan Ding, Nicholas Botzer, Tim Weninger

Modern deep learning systems like PyTorch and Tensorflow are able to train enormous models with billions (or trillions) of parameters on a distributed infrastructure. These systems require that the internal nodes have th…

GPUimage-classificationImage ClassificationLanguage Modeling+2

BlazeAIoT: A Modular Multi-Layer Platform for Real-Time Distributed Robotics Across Edge, Fog, and Cloud Infrastructures

2026-01-09 · Cedric Melancon, Julien Gascon-Samson, Maarouf Saad, Kuljeet Kaur 외 arxiv

The increasing complexity of distributed robotics has driven the need for platforms that seamlessly integrate edge, fog, and cloud computing layers while meeting strict real-time constraints. This paper introduces BlazeA…

AI-Native Network Controller: A Modular Framework for Safe Agentic Control of Multi-Domain Network Infrastructure

2026-04-20 · Merim Dzaferagic arxiv

The convergence of multiple network domains, including radio access, optical transport, and core networks, under unified intelligent control is a fundamental requirement for future 6G systems. This is important because e…

A Modular Benchmarking Infrastructure for High-Performance and Reproducible Deep Learning

2019-01-29 · Tal Ben-Nun, Maciej Besta, Simon Huber, Alexandros Nikolaos Ziogas 외

We introduce Deep500: the first customizable benchmarking infrastructure that enables fair comparison of the plethora of deep learning frameworks, algorithms, libraries, and techniques. The key idea behind Deep500 is its…

BenchmarkingDeep LearningVocal Bursts Intensity Prediction

NFDIcore 2.0: A BFO-Compliant Ontology for Multi-Domain Research Infrastructures

2024-09-16 · Oleksandra Bruns, Tabea Tietz, Joerg Waitelonis, Etienne Posthumus 외

This paper presents NFDIcore 2.0, an ontology compliant with the Basic Formal Ontology (BFO) designed to represent the diverse research communities of the National Research Data Infrastructure (NFDI) in Germany. NFDIcore…