paper-with-me

홈 › Papers

Omnivore: An Optimizer for Multi-device Deep Learning on CPUs and GPUs

2016-06-14 · Stefan Hadjis, Ce Zhang, Ioannis Mitliagkas, Dan Iter, Christopher Ré

We study the factors affecting training time in multi-device deep learning systems. Given a specification of a convolutional neural network, our goal is to minimize the time to train this model on a cluster of commodity CPUs and GPUs. We first focus on the single-node setting and show that by using standard batching and data-parallel techniques, throughput can be improved by at least 5.5x over state-of-the-art systems on CPUs. This ensures an end-to-end training speed directly proportional to the throughput of a device regardless of its underlying hardware, allowing each node in the cluster to be treated as a black box. Our second contribution is a theoretical and empirical study of the tradeoffs affecting end-to-end training time in a multiple-device setting. We identify the degree of asynchronous parallelization as a key factor affecting both hardware and statistical efficiency. We see that asynchrony can be viewed as introducing a momentum term. Our results imply that tuning momentum is critical in asynchronous parallel configurations, and suggest that published results that have not been fully tuned might report suboptimal performance for some configurations. For our third contribution, we use our novel understanding of the interaction between system and optimization dynamics to provide an efficient hyperparameter optimizer. Our optimizer involves a predictive model for the total time to convergence and selects an allocation of resources to minimize that time. We demonstrate that the most popular distributed deep learning systems fall within our tradeoff space, but do not optimize within the space. By doing this optimization, our prototype runs 1.9x to 12x faster than the fastest state-of-the-art systems.

📄 PDF Abstract BibTeX arXiv:1606.04487

Code (1)

HazyResearch/Omnivore 공식 구현

Similar Papers 제목 키워드 기반

Hybrid CPU-GPU Framework for Network Motifs

2016-08-18 · Ryan A. Rossi, Rong Zhou

Massively parallel architectures such as the GPU are becoming increasingly important due to the recent proliferation of data. In this paper, we propose a key class of hybrid parallel graphlet algorithms that leverages mu…

CPUGPU

Enabling On-Device Smartphone GPU based Training: Lessons Learned

2022-02-21 · Anish Das, Young D. Kwon, Jagmohan Chauhan, Cecilia Mascolo

Deep Learning (DL) has shown impressive performance in many mobile applications. Most existing works have focused on reducing the computational and resource overheads of running Deep Neural Networks (DNN) inference on re…

CPUGPU

Deep Optimizer States: Towards Scalable Training of Transformer Models Using Interleaved Offloading

2024-10-26 · Avinash Maurya, Jie Ye, M. Mustafa Rafique, Franck Cappello 외

Transformers and large language models~(LLMs) have seen rapid adoption in all domains. Their sizes have exploded to hundreds of billions of parameters and keep increasing. Under these circumstances, the training of trans…

CPUGPU

Boosting performance of computer vision applications through embedded GPUs on the edge

2025-11-03 · Fabio Diniz Rossi arxiv

Computer vision applications, especially those using augmented reality technology, are becoming quite popular in mobile devices. However, this type of application is known as presenting significant demands regarding reso…

Challenging GPU Dominance: When CPUs Outperform for On-Device LLM Inference

2025-05-09 · Haolin Zhang, Jeff Huang

The common assumption in on-device AI is that GPUs, with their superior parallel processing, always provide the best performance for large language model (LLM) inference. In this work, we challenge this notion by empiric…

CPUGPULarge Language ModelQuantization