paper-with-me

홈 › Papers

Dynamic Control Flow in Large-Scale Machine Learning

2018-05-04 · Yuan Yu, Martín Abadi, Paul Barham, Eugene Brevdo, Mike Burrows, Andy Davis, Jeff Dean, Sanjay Ghemawat, Tim Harley, Peter Hawkins, Michael Isard, Manjunath Kudlur, Rajat Monga, Derek Murray, Xiaoqiang Zheng

Many recent machine learning models rely on fine-grained dynamic control flow for training and inference. In particular, models based on recurrent neural networks and on reinforcement learning depend on recurrence relations, data-dependent conditional execution, and other features that call for dynamic control flow. These applications benefit from the ability to make rapid control-flow decisions across a set of computing devices in a distributed system. For performance, scalability, and expressiveness, a machine learning system must support dynamic control flow in distributed and heterogeneous environments. This paper presents a programming model for distributed machine learning that supports dynamic control flow. We describe the design of the programming model, and its implementation in TensorFlow, a distributed machine learning system. Our approach extends the use of dataflow graphs to represent machine learning models, offering several distinctive features. First, the branches of conditionals and bodies of loops can be partitioned across many machines to run on a set of heterogeneous devices, including CPUs, GPUs, and custom ASICs. Second, programs written in our model support automatic differentiation and distributed gradient computations, which are necessary for training machine learning models that use control flow. Third, our choice of non-strict semantics enables multiple loop iterations to execute in parallel across machines, and to overlap compute and I/O operations. We have done our work in the context of TensorFlow, and it has been used extensively in research and production. We evaluate it using several real-world applications, and demonstrate its performance and scalability.

📄 PDF Abstract BibTeX arXiv:1805.01772

Code (0)

등록된 구현이 없습니다.

Tasks

BIG-bench Machine LearningReinforcement Learning

Similar Papers 제목 키워드 기반

WAKESET: A Large-Scale, High-Reynolds Number Flow Dataset for Machine Learning of Turbulent Wake Dynamics

2026-02-01 · Zachary Cooper-Baldock, Paulo E. Santos, Russell S. A. Brinkworth, Karl Sammut arxiv

Machine learning (ML) offers transformative potential for computational fluid dynamics (CFD), promising to accelerate simulations, improve turbulence modelling, and enable real-time flow prediction and control-capabiliti…

m4: A Learned Flow-level Network Simulator

2025-03-03 · Chenning Li, Anton A. Zabreyko, Arash Nasr-Esfahany, Kevin Zhao 외

Flow-level simulation is widely used to model large-scale data center networks due to its scalability. Unlike packet-level simulators that model individual packets, flow-level simulators abstract traffic as continuous fl…

Modular Resource Centric Learning for Workflow Performance Prediction

2017-11-15 · Alok Singh, Mai Nguyen, Shweta Purawat, Daniel Crawl 외

Workflows provide an expressive programming model for fine-grained control of large-scale applications in distributed computing environments. Accurate estimates of complex workflow execution metrics on large-scale machin…

BIG-bench Machine LearningDistributed ComputingPredictionScheduling

Machine Learning-driven Multiscale MD Workflows: The Mini-MuMMI Experience

2025-07-10 · Loïc Pottier, Konstantia Georgouli, Timothy S. Carpenter, Fikret Aydin 외 arxiv

Computational models have become one of the prevalent methods to model complex phenomena. To accurately model complex interactions, such as detailed biomolecular interactions, scientists often rely on multiscale models c…

Learning Sampled-data Control for Swarms via MeanFlow

2026-03-20 · Anqi Dong, Yongxin Chen, Karl H. Johansson, Johan Karlsson arxiv

Steering large-scale swarms with only limited control updates is often needed due to communication or computational constraints, yet most learning-based approaches do not account for this and instead model instantaneous …

Decision Making