paper-with-me

Papers

Benchmarking Resource Usage for Efficient Distributed Deep Learning

2022-01-28 · Nathan C. Frey, Baolin Li, Joseph McDonald, Dan Zhao, Michael Jones, David Bestor, Devesh Tiwari, Vijay Gadepally, Siddharth Samsi

Deep learning (DL) workflows demand an ever-increasing budget of compute and energy in order to achieve outsized gains. Neural architecture searches, hyperparameter sweeps, and rapid prototyping consume immense resources that can prevent resource-constrained researchers from experimenting with large models and carry considerable environmental impact. As such, it becomes essential to understand how different deep neural networks (DNNs) and training leverage increasing compute and energy resources -- especially specialized computationally-intensive models across different domains and applications. In this paper, we conduct over 3,400 experiments training an array of deep networks representing various domains/tasks -- natural language processing, computer vision, and chemistry -- on up to 424 graphics processing units (GPUs). During training, our experiments systematically vary compute resource characteristics and energy-saving mechanisms such as power utilization and GPU clock rate limits to capture and illustrate the different trade-offs and scaling behaviors each representative model exhibits under various resource and energy-constrained regimes. We fit power law models that describe how training time scales with available compute resources and energy constraints. We anticipate that these findings will help inform and guide high-performance computing providers in optimizing resource utilization, by selectively reducing energy consumption for different deep learning tasks/workflows with minimal impact on training.

📄 PDF Abstract BibTeX arXiv:2201.12423

Code (0)

등록된 구현이 없습니다.

Tasks

BenchmarkingDeep LearningGPU

Similar Papers 제목 키워드 기반

Benchmarking Dynamic SLO Compliance in Distributed Computing Continuum Systems

2025-03-05 · Alfreds Lapkovskis, Boris Sedlak, Sindri Magnússon, Schahram Dustdar 외

Ensuring Service Level Objectives (SLOs) in large-scale architectures, such as Distributed Computing Continuum Systems (DCCS), is challenging due to their heterogeneous nature and varying service requirements across diff…

BenchmarkingCPUDistributed Computing

Applications in CityLearn Gym Environment for Multi-Objective Control Benchmarking in Grid-Interactive Buildings and Districts

2024-08-27 · Kingsley Nweye, Zoltan Nagy

It is challenging to coordinate multiple distributed energy resources in a single or multiple buildings to ensure efficient and flexible operation. Advanced control algorithms such as model predictive control and reinfor…

BenchmarkingModel Predictive Control

ResBench: Benchmarking LLM-Generated FPGA Designs with Resource Awareness

2025-03-11 · Ce Guo, Tong Zhao

Field-Programmable Gate Arrays (FPGAs) are widely used in modern hardware design, yet writing Hardware Description Language (HDL) code for FPGA implementation remains a complex and time-consuming task. Large Language Mod…

BenchmarkingCode GenerationDiversity

Design, Benchmarking and Explainability Analysis of a Game-Theoretic Framework towards Energy Efficiency in Smart Infrastructure

2019-10-16 · Ioannis C. Konstantakopoulos, Hari Prasanna Das, Andrew R. Barkan, Shiying He 외

In this paper, we propose a gamification approach as a novel framework for smart building infrastructure with the goal of motivating human occupants to reconsider personal energy usage and to have positive effects on the…

BenchmarkingDecision Making

EnviroLLM: Resource Tracking and Optimization for Local AI

2025-12-12 · Troy Allen arxiv

Large language models (LLMs) are increasingly deployed locally for privacy and accessibility, yet users lack tools to measure their resource usage, environmental impact, and efficiency metrics. This paper presents Enviro…