paper-with-me

홈 › Papers

Tune As You Scale: Hyperparameter Optimization For Compute Efficient Training

2023-06-13 · Abraham J. Fetterman, Ellie Kitanidis, Joshua Albrecht, Zachary Polizzi, Bryden Fogelman, Maksis Knutins, Bartosz Wróblewski, James B. Simon, Kanjun Qiu

Hyperparameter tuning of deep learning models can lead to order-of-magnitude performance gains for the same amount of compute. Despite this, systematic tuning is uncommon, particularly for large models, which are expensive to evaluate and tend to have many hyperparameters, necessitating difficult judgment calls about tradeoffs, budgets, and search bounds. To address these issues and propose a practical method for robustly tuning large models, we present Cost-Aware Pareto Region Bayesian Search (CARBS), a Bayesian optimization algorithm that performs local search around the performance-cost Pareto frontier. CARBS does well even in unbounded search spaces with many hyperparameters, learns scaling relationships so that it can tune models even as they are scaled up, and automates much of the "black magic" of tuning. Among our results, we effectively solve the entire ProcGen benchmark just by tuning a simple baseline (PPO, as provided in the original ProcGen paper). We also reproduce the model size vs. training tokens scaling result from the Chinchilla project (Hoffmann et al. 2022), while simultaneously discovering scaling laws for every other hyperparameter, via an easy automated process that uses significantly less compute and is applicable to any deep learning problem (not just language models).

📄 PDF Abstract BibTeX arXiv:2306.08055

Code (0)

등록된 구현이 없습니다.

Tasks

Bayesian OptimizationHyperparameter Optimization

Methods 이 논문이 사용한 방법론

Chinchilla 설명 없음

Similar Papers 제목 키워드 기반

A New Linear Scaling Rule for Private Adaptive Hyperparameter Optimization

2022-12-08 · Ashwinee Panda, Xinyu Tang, Saeed Mahloujifar, Vikash Sehwag 외

An open problem in differentially private deep learning is hyperparameter optimization (HPO). DP-SGD introduces new hyperparameters and complicates existing ones, forcing researchers to painstakingly tune hyperparameters…

Hyperparameter OptimizationImage Classification

Optimizing Millions of Hyperparameters by Implicit Differentiation

2019-11-06 · Jonathan Lorraine, Paul Vicol, David Duvenaud

We propose an algorithm for inexpensive gradient-based hyperparameter optimization that combines the implicit function theorem (IFT) with efficient inverse Hessian approximations. We present results about the relationshi…

Data AugmentationHyperparameter Optimization

When Losses Align: Gradient-Based Composite Loss Weighting for Efficient Pretraining

2026-05-08 · Ivan Karpukhin, Andrey Savchenko arxiv

Modern deep models are often pretrained on large-scale data with missing labels using composite objectives, where the relative weights of multiple loss terms act as hyperparameters. Tuning these weights with random searc…

Hyperparameter optimization of data-driven AI models on HPC systems

2022-03-02 · Eric Wulff, Maria Girone, Joosep Pata

In the European Center of Excellence in Exascale computing "Research on AI- and Simulation-Based Engineering at Exascale" (CoE RAISE), researchers develop novel, scalable AI technologies towards Exascale. This work exerc…

Bayesian OptimizationGraph Neural NetworkHyperparameter Optimization

Self-Tuning Networks: Bilevel Optimization of Hyperparameters using Structured Best-Response Functions

2019-03-07 · ICLR 2019 5 · Matthew MacKay, Paul Vicol, Jon Lorraine, David Duvenaud 외

Hyperparameter optimization can be formulated as a bilevel optimization problem, where the optimal parameters on the training set depend on the hyperparameters. We aim to adapt regularization hyperparameters for neural n…

Bilevel OptimizationData AugmentationHyperparameter Optimization