paper-with-me

홈 › Papers

Analysis of DAWNBench, a Time-to-Accuracy Machine Learning Performance Benchmark

2018-06-04 · Cody Coleman, Daniel Kang, Deepak Narayanan, Luigi Nardi, Tian Zhao, Jian Zhang, Peter Bailis, Kunle Olukotun, Chris Re, Matei Zaharia

Researchers have proposed hardware, software, and algorithmic optimizations to improve the computational performance of deep learning. While some of these optimizations perform the same operations faster (e.g., increasing GPU clock speed), many others modify the semantics of the training procedure (e.g., reduced precision), and can impact the final model's accuracy on unseen data. Due to a lack of standard evaluation criteria that considers these trade-offs, it is difficult to directly compare these optimizations. To address this problem, we recently introduced DAWNBench, a benchmark competition focused on end-to-end training time to achieve near-state-of-the-art accuracy on an unseen dataset---a combined metric called time-to-accuracy (TTA). In this work, we analyze the entries from DAWNBench, which received optimized submissions from multiple industrial groups, to investigate the behavior of TTA as a metric as well as trends in the best-performing entries. We show that TTA has a low coefficient of variation and that models optimized for TTA generalize nearly as well as those trained using standard methods. Additionally, even though DAWNBench entries were able to train ImageNet models in under 3 minutes, we find they still underutilize hardware capabilities such as Tensor Cores. Furthermore, we find that distributed entries can spend more than half of their time on communication. We show similar findings with entries to the MLPERF v0.5 benchmark.

📄 PDF Abstract BibTeX arXiv:1806.01427

Code (0)

등록된 구현이 없습니다.

Tasks

BenchmarkingBIG-bench Machine LearningGPU

Similar Papers 제목 키워드 기반

FastFusionNet: New State-of-the-Art for DAWNBench SQuAD

2019-02-28 · Felix Wu, Boyi Li, Lequn Wang, Ni Lao 외

In this technical report, we introduce FastFusionNet, an efficient variant of FusionNet [12]. FusionNet is a high performing reading comprehension architecture, which was designed primarily for maximum retrieval accuracy…

Reading ComprehensionRetrieval

Super-Convergence: Very Fast Training of Neural Networks Using Large Learning Rates

2017-08-23 · Leslie N. Smith, Nicholay Topin

In this paper, we describe a phenomenon, which we named "super-convergence", where neural networks can be trained an order of magnitude faster than with standard training methods. The existence of super-convergence is re…

Demystifying the MLPerf Benchmark Suite

2019-08-24 · Snehil Verma, Qinzhe Wu, Bagus Hanindhito, Gunjan Jha 외

MLPerf, an emerging machine learning benchmark suite strives to cover a broad range of applications of machine learning. We present a study on its characteristics and how the MLPerf benchmarks differ from some of the pre…

BIG-bench Machine LearningCPUGPUScheduling

Towards Scalable Distributed Training of Deep Learning on Public Cloud Clusters

2020-10-20 · Shaohuai Shi, Xianhao Zhou, Shutao Song, Xingyao Wang 외

Distributed training techniques have been widely deployed in large-scale deep neural networks (DNNs) training on dense-GPU clusters. However, on public cloud clusters, due to the moderate inter-connection bandwidth betwe…

GPU

Comparative Study of Machine Learning Models and BERT on SQuAD

2020-05-22 · Devshree Patel, Param Raval, Ratnam Parikh, Yesha Shastri

This study aims to provide a comparative analysis of performance of certain models popular in machine learning and the BERT model on the Stanford Question Answering Dataset (SQuAD). The analysis shows that the BERT model…

BIG-bench Machine LearningQuestion Answering