Probabilistic Rollouts for Learning Curve Extrapolation Across Hyperparameter Settings
We propose probabilistic models that can extrapolate learning curves of iterative machine learning algorithms, such as stochastic gradient descent for training deep networks, based on training data with variable-length learning curves. We study instantiations of this framework based on random forests and Bayesian recurrent neural networks. Our experiments show that these models yield better predictions than state-of-the-art models from the hyperparameter optimization literature when extrapolating the performance of neural networks trained with different hyperparameter settings.
Code (1)
Tasks
BIG-bench Machine LearningHyperparameter OptimizationSimilar Papers 제목 키워드 기반
Tune My Adam, Please!
The Adam optimizer remains one of the most widely used optimizers in deep learning, and effectively tuning its hyperparameters is key to optimizing performance. However, tuning can be tedious and costly. Freeze-thaw Baye…
Hyperparameter OptimizationArchitecture-Aware Learning Curve Extrapolation via Graph Ordinary Differential Equation
Learning curve extrapolation predicts neural network performance from early training epochs and has been applied to accelerate AutoML, facilitating hyperparameter tuning and neural architecture search. However, existing …
AutoMLNeural Architecture SearchHAMLET -- A Learning Curve-Enabled Multi-Armed Bandit for Algorithm Selection
Automated algorithm selection and hyperparameter tuning facilitates the application of machine learning. Traditional multi-armed bandit strategies look to the history of observed rewards to identify the most promising ar…
BIG-bench Machine LearningDEEP-BO for Hyperparameter Optimization of Deep Networks
The performance of deep neural networks (DNN) is very sensitive to the particular choice of hyper-parameters. To make it worse, the shape of the learning curve can be significantly affected when a technique like batchnor…
Bayesian OptimizationHyperparameter OptimizationLight curve completion and forecasting using fast and scalable Gaussian processes (MuyGPs)
Temporal variations of apparent magnitude, called light curves, are observational statistics of interest captured by telescopes over long periods of time. Light curves afford the exploration of Space Domain Awareness (SD…
Gaussian ProcessesPose EstimationTime SeriesTime Series Analysis